DevOps Engineer
Designs, builds, and operates reliable cloud infrastructure for real-time voice AI systems. Owns Kubernetes clusters, CI/CD pipelines, observability, and security using AWS and IaC tools. Requires 5+ years DevOps experience with strong Python and async programming skills.
About the job
Responsibilities
- Design, build, and operate highly reliable cloud infrastructure that powers real-time voice AI systems with extremely low latency and high availability.
- Own Kubernetes clusters end-to-end: provisioning, scaling, upgrades, networking, and debugging production incidents under real customer load.
- Build, maintain, and evolve infrastructure as code using tools like Terraform, Pulumi, or CloudFormation to ensure repeatable, auditable, and secure environments across staging and production.
- Create and operate CI/CD pipelines that enable fast, safe iteration across multiple microservices and teams.
- Design and maintain observability systems (metrics, logs, traces, alerting) to detect failures early and rapidly diagnose production issues.
- Partner with backend engineers to translate application requirements into scalable, secure infrastructure and clean deployment workflows.
- Harden systems through strong security practices including IAM, secrets management, network isolation, and least-privilege access controls.
- Optimize cloud performance and costs while maintaining reliability, developer velocity, and customer experience.
- Implement and operate GitOps-driven deployment workflows, using Git as the source of truth for infrastructure and application state, enabling safe, auditable, and automated rollouts.
- Lead incident response: investigate outages, coordinate fixes, write postmortems, and drive systemic reliability improvements.
- Continuously improve resilience through load testing, chaos testing, capacity planning, and proactive infrastructure upgrades.
Qualifications
- 5+ years as a DevOps engineer
- Experience writing async web apps using FastAPI in Python
- Builder of APIs, Clouds, CI/CD pipelines
- Experience with IaC, AWS, Database Management at scale
- Understanding of good architecture, security practices
- Strong technical and communication skills
- Extensive experience with AWS & Kubernetes
Software Stack
- Backend: Python, microservices, async programming
- Cloud & Infrastructure: AWS, GCP, Kubernetes, Redis, ArgoCD, GitOps
- Databases: Firebase, Supabase (PostgreSQL)
- Frontend: Next.js
- Observability & Monitoring: Datadog, logging, metrics, tracing
- Telephony & Voice AI: SIP, voice APIs, real-time call handling
- Other tools & practices: CI/CD, automated testing, resilient architecture
Skills
Kubernetes, AWS, Terraform, CI/CD, GitOps, Python, FastAPI, Argo CD, Datadog, Redis, Postgres, GCP, Pulumi, CloudFormation, IAM
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.