Skip to content
OpenAIOpenAI

Software Engineer, Frontier Systems

Builds infrastructure to monitor, detect, remediate, and verify hardware health across global GPU/CPU clusters at hyperscale. Owns node lifecycle workflows and partners with teams to ensure compute reliability for AI training and inference. Requires 7+ years experience with Python, distributed systems, and operational tooling.

About the job

Responsibilities

  • Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure.
  • Build and evolve health checks that detect, remediate, and verify failures at scale.
  • Ensure critical health checks execute with minimal latency to maximize workload uptime.
  • Investigate hardware failures and system-level issues across large-scale compute environments.
  • Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes.
  • Build automation and tooling that enables global cluster management with minimal manual intervention.
  • Partner with workload, reliability, and provider teams to integrate health signals into training and inference systems.

Requirements

  • 7+ years of industry experience in software or infrastructure engineering.
  • Strong proficiency with Python and shell scripting.
  • Experience building large-scale distributed systems or infrastructure platforms.
  • Comfort digging into noisy operational data using SQL, PromQL, or similar tooling.
  • Experience building reproducible analyses and operational tooling.
  • Strong systems debugging and operational instincts with an ownership mindset.

Nice-to-Haves

  • Experience with low-level hardware systems and Linux tooling (e.g. PCIe, InfiniBand, RoCE, networking, power management, kernel performance tuning, FW/SW debugging).
  • Experience operating or debugging large-scale GPU or accelerator clusters.
  • Expertise in network operations, observability, or systems telemetry.
  • Experience with automated remediation systems or fleet lifecycle management.
  • Experience improving reliability, utilization, or workload uptime in distributed compute environments.

Skills

Python, SQL, Promql, Kubernetes, Linux, Pcie, InfiniBand, Roce, GPU, Distributed Systems

Vapi

Vapi

San Francisco, CA

Member of Technical Staff, Release Engineer
$235k+/yrHybrid7+ YOEDevOps / SRE

Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.

The Voleon Group

The Voleon Group

Berkeley, CA
Senior Software Engineer, Developer Experience
$225k+/yrHybrid5+ YOEDevOps / SRE

Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.

Descript

Descript

San Francisco, CA

Software Engineer, Infrastructure
$220k+/yrRemote8+ YOEDevOps / SRE

Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.

Zoox

Zoox

Foster City, CA

Senior Software Engineer - Pipeline Infrastructure & Integration
$219k+/yrHybrid7+ YOEDevOps / SRE

Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.

Skydio

Skydio

San Mateo, CA

Senior Software Engineer, Developer Productivity
$200k+/yrOn-site5+ YOEDevOps / SRE

Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.