Skip to content
DescriptDescript

Software Engineer, Infrastructure

Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.

About the job

Responsibilities

  • Own the platform foundation, including Google Cloud, Kubernetes, Temporal, GPU infrastructure, deployment, and rollback systems.
  • Participate in the on-call rotation and improve incident response, reliability, and actionability.
  • Build and operate AI infrastructure for GPU capacity, model training and inference pipelines, and production serving.
  • Manage infrastructure cost modeling, metering, attribution, and cost-performance trade-offs.
  • Own security boundaries including identity and access management, secrets, least privilege, and software supply-chain integrity.
  • Develop infrastructure-as-code, runbooks, observability, tests, standards, and release practices.
  • Partner with engineering teams to understand needs, set the platform roadmap, and make architecture decisions.
  • Provide architectural direction, code reviews, mentoring, and clear technical documentation.

Requirements

  • 8+ years building and operating production distributed systems or server-side systems with a strong infrastructure focus.
  • Experience operating production systems where reliability and failure consequences matter.
  • Experience carrying a pager, commanding incidents, and performing rollbacks.
  • Practical use of SLOs and error budgets.
  • Production experience with a major cloud provider, Kubernetes, and infrastructure-as-code.
  • Experience owning an architecture or migration from planning through launch.
  • Ability to independently identify, scope, gain support for, and deliver unowned work.
  • Strong debugging skills, including creating minimal reproductions, analyzing logs, and writing targeted checks.
  • Ability to use AI agents effectively and critically in engineering work.

Nice-to-haves

  • GPU and machine-learning infrastructure, including capacity planning, training or inference pipelines, and model-serving cost and latency optimization.
  • Production security engineering, including IAM, secrets management, software supply-chain security, and least privilege.
  • Cloud cost modeling, commitment strategies, reservations, and unit economics.
  • CI/CD at monorepo scale and developer-environment tooling.
  • Media, video, or GPU-backed workload experience.
  • Experience on small teams responsible for broad infrastructure surfaces.

Compensation and Benefits

  • Base salary: $220,000–$292,000, plus equity and benefits.
  • Healthcare package, 401(k) matching, catered lunches, and flexible vacation time.

Skills

GCP, Kubernetes, Temporal, Gpu Infrastructure, Machine Learning Infrastructure, CI/CD, Infrastructure As Code, Distributed Systems, SLOs, Error Budgets, IAM, Secrets Management, Observability, Monorepos, Cloud Cost Modeling

Zoox

Zoox

Foster City, CA

Senior Software Engineer - Pipeline Infrastructure & Integration
$219k+/yrHybrid7+ YOEDevOps / SRE

Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.

The Voleon Group

The Voleon Group

Berkeley, CA
Senior Software Engineer, Developer Experience
$225k+/yrHybrid5+ YOEDevOps / SRE

Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.

Vapi

Vapi

San Francisco, CA

Member of Technical Staff, Release Engineer
$235k+/yrHybrid7+ YOEDevOps / SRE

Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.

Skydio

Skydio

San Mateo, CA

Senior Software Engineer, Developer Productivity
$200k+/yrOn-site5+ YOEDevOps / SRE

Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.

Anyscale

Anyscale

San Francisco, CA

Senior Site Reliability Engineer, Platform Infrastructure
$200k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.