Site Reliability / Infrastructure Engineer
Own reliability and infrastructure for large-scale video and social systems, including on-call response, postmortems, database scaling, infrastructure as code, and CI/CD. The role requires deep GCP, Kubernetes, Terraform, Elasticsearch, and production database experience.
About the job
Responsibilities
- Own the infrastructure on-call rotation and respond to production incidents.
- Drive incident response, communicate clearly during P0 incidents, and lead actionable postmortems.
- Scale and maintain infrastructure supporting billions of clips, video ingestion pipelines, and social features.
- Partner directly with engineering teams to meet infrastructure needs.
- Scale and shard relational databases in production.
- Build and maintain infrastructure as code and CI/CD systems.
Requirements
- Strong proficiency with Terraform and experience owning infrastructure as code at scale.
- Hands-on experience running Elasticsearch for user-facing features.
- Deep experience with Google Cloud Platform, including Kubernetes, VPC, IAM, Cloud Logging, and managed services.
- Hands-on production experience scaling and sharding MySQL or PostgreSQL databases.
- Production incident response experience, including P0 incidents and postmortems.
- Experience with GitHub Actions in a production environment.
- Experience working at startups or scale-ups in rapidly growing environments.
- Strong communication, judgment, and ability to distinguish durable fixes from temporary patches.
Technology Stack
- Google Cloud Platform
- Kubernetes
- Terraform
- Elasticsearch
- MySQL
- PostgreSQL
- GitHub Actions
- Salt
- CircleCI
- Redis
- RabbitMQ
Skills
Terraform, Elasticsearch, GCP, Kubernetes, Vpc, IAM, Cloud Logging, MySQL, Postgres, GitHub Actions, Salt, CircleCI, Redis, RabbitMQ
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.