Site Reliability Engineer
Owns production infrastructure for clinical AI platform, ensuring 99.9%+ stability. Designs/scales Kubernetes-based systems, optimizes TypeScript/Python/ML CI/CD pipelines, and manages Terraform IaC in high-velocity environment.
About the job
Responsibilities
- Own the entire production environment and improve the development experience.
- Design, implement, and maintain the production environment, having previously handled 500+ machine deployments.
- Own containerized infrastructure, leveraging deep expertise in Kubernetes and Helm to manage deployment, scaling, and operational health.
- Optimize and streamline both TypeScript and Python/ML deployment pipelines to support high-velocity feature release while maintaining highest reliability.
- Support Developer Experience (DevX) work to streamline developer workflows, enhance tool proficiency, and improve CI/CD systems.
- Manage and maintain infrastructure definitions using Terraform.
Requirements
- Deep, demonstrable experience with Kubernetes, Helm, and Terraform.
- Proven ability to architect and maintain complex, distributed systems with high-availability requirements.
- Hands-on experience optimizing deployment pipelines for both application code (TypeScript) and machine learning models (Python/ML).
- Experience with PostgreSQL, Redis, Kafka.
- Excitement about working five days per week in San Francisco office.
- Intensity and technical mastery to own mission-critical infrastructure; thrive on owning complex systems, scaling deployments, automating, and problem-solving.
Skills
Kubernetes, Helm, Terraform, TypeScript, Python, CI/CD, Postgres, Redis, Kafka, Infrastructure As Code
Similar jobs
DevOps / SRE jobsBuild and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.