Build, scale, and operate OpenAI's global compute infrastructure for frontier AI models like GPT-5.6. Solve complex cross-disciplinary problems spanning distributed systems, hardware, ML infrastructure, power/cooling, manufacturing, supply chain, and data center development at unprecedented scale.
230k – 490k/yr
On-site5+ YOEDevOps / SRE
About the role
Key Responsibilities
Help build, scale, and operate OpenAI’s global compute infrastructure.
Solve complex problems across software, hardware, manufacturing supply chain, and data center systems.
Improve the reliability, performance, efficiency, and scalability of critical infrastructure.
Partner with cross-functional teams to bring new compute capacity online quickly and reliably.
Identify bottlenecks across technical, operational, and physical systems, and develop practical solutions.
Build tools, processes, systems, or infrastructure that improve execution at scale.
Contribute to the long-term architecture and operational maturity of OpenAI’s compute footprint.
Qualifications
Experience building, scaling, or operating complex technical systems.
Enjoy working on ambiguous, high-impact problems where the path forward is not always defined.
Comfortable collaborating across disciplines, including software, hardware, operations, and physical infrastructure.
Strong technical judgment and a bias toward execution.
Care deeply about reliability, speed, safety, and operational excellence.
Excited by the challenge of building infrastructure at unprecedented scale.
Work directly supports the development and deployment of frontier AI.
Preferred Skills
Experience with AI infrastructure, high-performance computing, distributed systems, GPU clusters, or cloud-scale platforms.
Worked on hardware systems, manufacturing, supply chain, data center development, or large capital infrastructure projects.
Domain expertise in civil, controls, mechanical, hardware, electrical, thermal, power, networking, or facilities engineering.
Helped bring new technical platforms, data centers, factories, or large-scale systems from concept to production.
Experience operating in fast-moving environments where technical depth and execution speed both matter.
Builds and improves CI/CD, testing, validation, and release tooling for OpenAI's inference runtime teams to ensure reliable, performant model deployments across ChatGPT, API, and research workloads. Requires strong Python skills, developer productivity experience, and high ownership in ambiguous environments.
230k – 385k/yr
On-siteDevOps / SRE
Software Engineer, Core Network Engineering
OpenAISan Francisco, CA
Builds and operates high-performance networking infrastructure for OpenAI's large-scale AI training and inference, focusing on host networking, datacenter fabrics, and WAN systems. Optimizes latency, reliability, and scalability using technologies like RDMA, InfiniBand, and RoCE; requires strong systems programming in C++, Python, or Go.
230k – 342k/yr
On-siteDevOps / SRE
Software Engineer, Productivity - Model Performance
OpenAISan Francisco, CA
Builds and improves developer tools, CI/CD pipelines, and testing workflows to boost productivity for OpenAI's model performance engineering teams. Requires strong Python skills, experience with developer infrastructure, and ability to work in ambiguous environments.
230k – 385k/yr
On-siteDevOps / SRE
Software Engineer, Productivity - Networking
OpenAISan Francisco, CA
Enhances developer productivity for OpenAI's networking team by improving build systems, CI/CD pipelines, test harnesses, and workflows for C++ and Python codebases in multi-server environments. Requires experience with developer tools and infrastructure automation.
230k – 385k/yr
On-siteDevOps / SRE
Software Engineer, Compute Infrastructure
OpenAISan Francisco, CA +2
Builds and optimizes large-scale compute infrastructure for AI workloads, spanning hardware automation, distributed systems, Kubernetes orchestration, networking, storage, and developer tools. Requires strong systems engineering experience in performance, reliability, and production infrastructure.