Software Engineer, Frontier Clusters Infrastructure
Builds and scales massive Kubernetes clusters for OpenAI's frontier supercomputers, automates bare-metal provisioning, and ensures reliability across data centers for AI model training. Requires expertise in distributed systems, Kubernetes operations, and infrastructure automation.
About the job
Responsibilities
- Spin up and scale large Kubernetes clusters, including automation for provisioning, bootstrapping, and cluster lifecycle management
- Build software abstractions that unify multiple clusters and present a seamless interface to training workloads
- Own node bring-up from bare metal through firmware upgrades, ensuring fast, repeatable deployment at massive scale
- Improve operational metrics such as reducing cluster restart times (e.g., from hours to minutes) and accelerating firmware or OS upgrade cycles
- Integrate networking and hardware health systems to deliver end-to-end reliability across servers, switches, and data center infrastructure
- Develop monitoring and observability systems to detect issues early and keep clusters stable under extreme load
Requirements
- Deep experience operating or scaling Kubernetes clusters or similar container orchestration systems in high-growth or hyperscale environments
- Strong programming or scripting skills (Python, Go, or similar) and familiarity with Infrastructure-as-Code tools such as Terraform or CloudFormation
- Comfortable with bare-metal Linux environments, GPU hardware, and large-scale networking
- Experience as an infrastructure, systems, or distributed systems engineer in large-scale or high-availability environments
- Strong knowledge of Kubernetes internals, cluster scaling patterns, and containerized workloads
- Proficiency in cloud infrastructure concepts (compute, networking, storage, security) and in automating cluster or data center operations
Nice-to-Haves
- Background with GPU workloads, firmware management, or high-performance computing
Skills
Kubernetes, Python, Go, Terraform, Linux, GPU, Distributed Systems, Infrastructure As Code, Bare Metal, Networking
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.