Reliability Engineer
Designs and maintains reliable infrastructure for petabyte-scale video processing across multi-cloud environments. Owns incident response, observability (Prometheus, OpenTelemetry), security, and CI/CD tooling with 3+ years experience in scalable systems.
About the job
What You’ll Do
- Work with engineering to design and validate the infrastructure powering PB-scale workloads
- Build and maintain Terraform-managed multi-cloud deployments
- Improve cloud and data security (SSO, IAM, least privilege, auditability)
- Own incident response and harden systems against failure
- Develop CI/CD systems that minimize user error and maximize safety
- Build monitoring + alerting platforms (Prometheus, OpenTelemetry, VictoriaMetrics)
- Wrap internal reliability tooling with simple UIs for engineers
Requirements
- 3+ years building internal infrastructure at scale
- Experience on-call for Sev 0 / Sev 1 production incidents (L3 preferred)
- Strong cloud experience (GCP, AWS, Oracle, Cloudflare, etc.)
- Deep Infrastructure-as-Code experience (Terraform preferred)
- Familiarity with Argo, Helm, Kustomize, or similar deployment tools
- Experience operating observability systems (Prometheus, OTel, VictoriaMetrics)
- Backend fundamentals in Python, Go, Rust, or C++
- Strong networking + security intuition, including SSO implementation
- High ownership mindset over critical systems
Bonus
- Experience building lightweight internal tooling (APIs, dashboards, Svelte)
- Familiarity with object storage systems (“buckets”)
- Active GitHub or portfolio projects
Benefits
- 401k + Full Health Insurance
- Breakfast, Lunch, and Dinner covered and your choice of snacks
- Ubers covered home
Skills
Terraform, GCP, AWS, Prometheus, OpenTelemetry, Victoriametrics, Argo, Helm, Kustomize, Python, Go, Rust, C++, Cloudflare, Svelte
Similar jobs
DevOps / SRE jobsBuild and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.
Leads on-site deployment of data center physical infrastructure, managing contractors, performing QA/QC on fiber optics and cabling, and ensuring compliance with standards. Requires 5+ years experience, SME-level fiber optic expertise, bachelor's degree, and 40% travel readiness.
The Python Engineer will improve and operate trading systems, support integrations with asset classes and prime brokers, and handle monitoring, incidents, and performance optimization. The role requires 3+ years of experience, strong Python and Linux skills, and familiarity with market data and order-entry systems.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
The DevOps Engineer will build and operate reliable infrastructure, deployment workflows, and observability for data pipelines and AI/ML systems. The role requires at least three years of DevOps, SRE, or infrastructure experience plus strong cloud, Terraform, containerization, and MLOps expertise.