Member of Technical Staff, DevOps
The Member of Technical Staff, DevOps will own progressive delivery, GitOps, and on-demand environment tooling to improve deployment safety and speed for engineering teams. This role requires a platform-as-a-product mindset and experience with infrastructure as code and CI/CD pipelines.
About the job
Why We’re Hiring This Role:
Three of our worst recent incidents — Nov 29 config rollout, Dec 23 duplicate messages, Oct 13 egress proxy — were resolved by rollback. That’s the SDLC gap this role closes. Vapi engineers are your users, and the deploy pipeline, preview environments, cell-creation tooling, and oncall tooling are products with SLAs, docs, and feedback loops.
You’ll own progressive delivery (canary, blue/green, automated rollback, soak periods), the GitOps story across multiple clusters and regions, and the on-demand environment tooling that’s on the Q3 roadmap. Success is measured by how fast every other team ships safely.
What You’ll Do:
- 30 Day: Get fluent in the Pulumi stacks, the ArgoCD setup, and GitHub Actions pipelines. Sit with engineers from agents and FDE teams to find the top 3 deploy pain points. Land a quality-of-life improvement to the deploy pipeline.
- 60 Day: Own progressive delivery end-to-end — canary, automated rollback, soak — for at least one critical service path. Ship the first version of cell-creation tooling or preview environments. Make the deploy pipeline measurably faster (lead time, MTTR for failed deploys).
- 90 Day: Roll out progressive delivery as the default across services. Establish SLAs and a feedback loop with engineering teams. Own the developer-platform roadmap and partner with Infra and SRE on cell creation, multi-region rollouts, and oncall tooling.
Who You Are:
Must-haves
- You have a platform-as-a-product mindset — you treat internal engineers as customers, with SLAs, docs, and feedback loops, not tickets and ad-hoc help.
- You’ve operated Pulumi (TypeScript) or Terraform at scale (40+ stacks, multi-region) and you’ve felt the pain when IaC sprawl gets ahead of you.
- You’ve run ArgoCD or equivalent GitOps for deploying applications across multiple clusters.
- You’ve built progressive delivery in production — canary, blue/green, automated rollback, soak periods. You can describe a real rollout that automated rollback caught.
- You’ve designed CI/CD pipelines (GitHub Actions preferred) for many services and Dockerfiles, not just one repo.
- You’ve built deploy tooling for on-demand environments — preview envs, dev deployments, or cell creation.
Nice-to-haves
- You’ve written Go for platform services (Vapi’s canary-manager is Go).
- You’ve operated developer platforms at a mid-stage infra-heavy company or a DevEx team at a larger shop.
Tech stack you’ll work in
- Languages: TypeScript (primary, for Pulumi and tooling), Go (for canary-manager and platform services), Bash.
- IaC: Pulumi (TypeScript) at scale (40+ stacks across regions), Terraform.
- GitOps and deploy: ArgoCD (multi-cluster), GitHub Actions, 15+ Dockerfiles.
- Progressive delivery: canary, blue/green, automated rollback, soak periods (canary-manager Go service).
- Orchestration: Kubernetes on EKS (multi-cluster, multi-region).
- Vapi services you’ll touch: canary-manager, cell-creation tooling, preview env tooling.
Where you likely come from
Vercel, Render, Railway, Fly, Temporal, Cockroach (mid-stage infra-heavy), or DevEx/Platform teams at Stripe, Shopify, Airbnb, or Block.
Weak fit: classic AWS sysadmin, or someone whose CI/CD experience is mostly Jenkins GUI-level.
Skills
Pulumi, Argo CD, GitHub Actions, TypeScript, Go, Bash, Terraform, Docker, Kubernetes, EKS
Similar jobs
DevOps / SRE jobsBuild and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.