Skip to content
VapiVapi

Member of Technical Staff, Infrastructure Engineer

Build software, automation, and observability for Vapi’s latency-sensitive distributed infrastructure. The role owns reliability improvements across incident response, capacity, performance, and failure prevention, requiring senior or staff-level experience with SRE and cloud infrastructure.

About the job

Responsibilities

  • Learn Vapi’s architecture, production environment, incident history, and reliability practices.
  • Build tooling and automation for distributed, latency-sensitive production systems.
  • Improve observability, incident response, capacity planning, performance, and production automation.
  • Reduce manual operational work and improve detection, diagnosis, and response to failures.
  • Own a reliability workstream and become the primary owner for a meaningful part of the reliability surface.
  • Deliver durable improvements to failure prevention and recovery.
  • Propose roadmaps for future reliability investments.
  • Collaborate with Infrastructure and product engineering teams.

Requirements

  • Senior- or staff-level software engineering experience.
  • Meaningful SRE, production engineering, or infrastructure experience with distributed systems.
  • Ability to write production-quality software and build reliability tooling or automation.
  • Deep experience with observability, incident response, failure analysis, capacity, and production reliability practices.
  • Experience with Kubernetes, networking, and cloud infrastructure.
  • Ability to debug across application and infrastructure boundaries.
  • Strong judgment around failure modes and balancing reliability investments with product and engineering velocity.

Nice-to-haves

  • Experience with real-time networking or telephony.
  • Experience with Envoy, Postgres, Redis, Kafka, Aurora, ClickHouse, or Google-style SRE practices.

Compensation and Benefits

  • Base salary of $280,000 to $314,000.
  • Equity ownership.
  • Medical, dental, and vision coverage.
  • Flexible time off.
  • Quarterly off-sites, catered meals, transportation, gym benefits, and a $10,000 annual learning and development budget.

Skills

SRE, Distributed Systems, Kubernetes, Networking, Cloud Infrastructure, Observability, Incident Response, Failure Analysis, Capacity Planning, Production Automation, Envoy, Postgres, Redis, Kafka, Aurora

Anthropic

Anthropic

Austin, TX
Data Center Operations Lead - Partner Site Operations
$320k+/yrHybrid8+ YOEDevOps / SRE

Leads operations outcomes for partner-operated data center sites, directing vendors, defining operational standards, and ensuring deployment velocity, availability, repair performance, and incident response. Requires 8+ years in data center or infrastructure operations, vendor oversight experience, and hands-on server, network, and rack-level expertise.

Vapi

Vapi

San Francisco, CA

Member of Technical Staff, Release Engineer
$235k+/yrHybrid7+ YOEDevOps / SRE

Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.

Descript

Descript

San Francisco, CA

Software Engineer, Infrastructure
$220k+/yrRemote8+ YOEDevOps / SRE

Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.

Zoox

Zoox

Foster City, CA

Senior Software Engineer - Pipeline Infrastructure & Integration
$219k+/yrHybrid7+ YOEDevOps / SRE

Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.

Tennr

Tennr

New York, NY

Senior Infrastructure Engineer
$200k+/yrOn-site5+ YOEDevOps / SRE

Own and scale Tennr’s AWS infrastructure, Kubernetes environments, and infrastructure-as-code foundation across development, staging, and production. The role requires 5–8 years of infrastructure, platform, or DevOps experience, strong Kubernetes and AWS expertise, and hands-on production ownership.