Staff Site Reliability Engineer
Zoox is seeking a Staff Site Reliability Engineer to lead source control, owning the technical strategy and roadmap for their Git-based monorepo. This role involves migrating from GitHub Enterprise to GitHub Cloud, building developer tooling, and partnering with various teams to enhance source control as a strategic asset.
About the job
Zoox is building fully autonomous vehicles designed from the ground up for robotaxi service — and as we scale toward launch, our engineering infrastructure has to scale with us. We're hiring a Staff Site Reliability Engineer to own source control at Zoox, serving as the technical lead for the platform that every engineer depends on daily. This is a hands-on lead role. You'll set the direction for our Git-based monorepo, drive a potential migration to GitHub Cloud, build the tooling that smooths developer workflows, and partner with platform, security, and product teams to make source control a strategic asset rather than a bottleneck. The ideal candidate leads by building — with deep expertise in monorepo management (branching strategies, code review workflows, access controls, CI/CD) and a track record of improving developer productivity through code, not just guidance. You take ownership, measure what matters, and have the technical credibility to make decisions stick. In this role, you will: Own the technical strategy and roadmap for source control — GitHub Enterprise today, GitHub Cloud tomorrow — and write the core code that makes it real. Plan and execute a full migration from GitHub Enterprise to GitHub Cloud: scoping, design, implementation, cutover, and validation, in partnership with security, platform, and engineering teams. Build and operate monorepo scalability improvements across repo structure, branch strategy, code ownership, large-file handling, and CI integration — including the scripts, hooks, and automation that make them stick. Develop the guardrails and developer-facing tooling that reduce toil for SRE and product engineering alike, measured by build times, merge friction, and interrupt volume. Set the technical bar through code reviews, design proposals, and pairing — and serve as the hands-on authority for source control decisions across Zoox engineering
Qualifications
- 5+ years operating GitHub Enterprise (or equivalent) at scale, with hands-on experience scaling a large monorepo for an org of several hundred or more engineers.
- Strong CI/CD integration background (Buildkite, GitHub Actions, Jenkins, or GitLab CI) across code review workflows, build triggers, branch policies, and merge automation.
- Experience with infrastructure as code (Terraform, Pulumi, or equivalent), major cloud platforms, and executing platform migrations against active codebases.
- Strong technical leadership and written communication; able to drive alignment across infrastructure, security, and product engineering.
Bonus Qualifications
- Direct experience migrating from GitHub Enterprise Server to GitHub Cloud.
- Familiarity with monorerepo build tooling (Bazel, Buck, or similar) and supply-chain hardening practices (signed commits, dependency governance, secret scanning).
- Experience with code review tooling layered on top of GitHub (Reviewable, Gerrit, or similar).
Skills
Github Enterprise, CI/CD, Buildkite, GitHub Actions, Jenkins, Gitlab Ci, Terraform, Pulumi, Bazel, Buck
Similar jobs
DevOps / SRE jobsStaff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.
Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.
Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.