Senior Production Engineer
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
About the job
Responsibilities
- Design, build, scale, secure, and maintain production services.
- Write or improve code frameworks for monitoring, logging, database access, and authentication.
- Build and optimize core infrastructure to support new features and functionality.
- Join the on-call rotation, support software engineers, and debug complex problems.
- Collaborate across technical and cross-functional teams.
Requirements
- 5+ years of production, site reliability, or DevOps engineering experience.
- Experience writing high-quality, well-tested code.
- Familiarity with cloud infrastructure and infrastructure-as-code solutions.
- Experience with AWS, Kubernetes, and Terraform.
- Experience programming in high-level languages such as Python is preferred.
- A computer science degree or equivalent experience is ideal but not required.
- Openness to relocate to Los Angeles, California, or ability to travel to the LA office monthly while working remotely.
Compensation
- Salary range: $200,000-$238,000 annually.
Skills
AWS, Kubernetes, Terraform, Python, Monitoring, Logging, Database Access, Authentication, Infrastructure As Code, Site Reliability
Similar jobs
DevOps / SRE jobsBuild and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.
Leads design, deployment, and operation of secure distributed cloud systems for public-sector and air-gapped environments. Requires active or obtainable TS/SCI clearance with polygraph, U.S. citizenship, and 7+ years of production experience.
Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.