Build and maintain CI/CD pipelines, orchestration, and automation for large-scale robotics simulation (SIL/HIL) to support model training, evaluation, and RL workloads at OpenAI. Requires strong infra, distributed systems, and Python/C++/Rust experience.
230k – 385k/yr
Hybrid5+ YOEDevOps / SRE
About the role
Responsibilities
Build and maintain presubmit checks, continuous integration and deployment pipelines for simulation code, environments, and tasks so simulation artifacts are testable, versioned, and reproducible.
Implement end-to-end automation to run model evaluation in sim (SIL) and orchestrate HIL runs; compute realism and task metrics, generate dashboards and alerts, and ensure evaluation is repeatable and auditable.
Create robust APIs and connectors so research, training, and data-collection systems can schedule, seed, and evaluate batches of simulations; support RL rollouts, imitation-data collection, and presubmit model checks.
Build scheduling, batching and orchestration for running very large numbers of concurrent rollouts (target tens of thousands of rollouts / large RL workloads), solve engine-level scaling (parallelization, batching multiple runs per engine), and optimize cloud/GPU runtime reliability.
Produce metrics and tooling for measuring simulation health, throughput, fidelity regressions, and cost; create presubmit / canary tests that catch sim regressions early.
Implement artifact versioning, environment immutability (images / asset versions), experiment provenance, and policies for resource quotas and cost control across the sim farm.
Work closely with Sim Environments, Sim Realism, research, and ops to close the loop—ensuring simulation improvements directly translate into better model evaluation and training results.
Requirements
Deep software engineering & infra experience: built CI/CD at scale, authored reliable pipelines, and shipped production services that coordinate many moving parts.
Comfortable with distributed compute and cloud GPU workloads: know how to get many sims running concurrently (scheduling, batching, GPU orchestration) and optimize throughput/cost.
Built or maintained HIL/SIL workflows or other sim↔hardware integrations and understand the operational challenges of bridging software and hardware testbeds.
Strong with automation, observability and metrics: enjoy instrumenting systems, defining meaningful KPIs, and surfacing regressions early.
Can design APIs and developer ergonomics so research and SWE teams can easily submit jobs, reproduce experiments, and interpret results.
Experience with Python/C++/Rust, container orchestration (Kubernetes), distributed task queues, and CI systems.
Enjoy collaborating across teams to turn experimental simulation work into dependable production tooling.
Nice-to-Haves
Worked with RL tooling, task generators, or large-scale data pipelines.
Build, scale, and operate OpenAI's global compute infrastructure for frontier AI models like GPT-5.6. Solve complex cross-disciplinary problems spanning distributed systems, hardware, ML infrastructure, power/cooling, manufacturing, supply chain, and data center development at unprecedented scale.
Builds and improves CI/CD, testing, validation, and release tooling for OpenAI's inference runtime teams to ensure reliable, performant model deployments across ChatGPT, API, and research workloads. Requires strong Python skills, developer productivity experience, and high ownership in ambiguous environments.
230k – 385k/yr
On-siteDevOps / SRE
Software Engineer, Core Network Engineering
OpenAISan Francisco, CA
Builds and operates high-performance networking infrastructure for OpenAI's large-scale AI training and inference, focusing on host networking, datacenter fabrics, and WAN systems. Optimizes latency, reliability, and scalability using technologies like RDMA, InfiniBand, and RoCE; requires strong systems programming in C++, Python, or Go.
230k – 342k/yr
On-siteDevOps / SRE
Software Engineer, Productivity - Model Performance
OpenAISan Francisco, CA
Builds and improves developer tools, CI/CD pipelines, and testing workflows to boost productivity for OpenAI's model performance engineering teams. Requires strong Python skills, experience with developer infrastructure, and ability to work in ambiguous environments.
230k – 385k/yr
On-siteDevOps / SRE
Software Engineer, Productivity - Networking
OpenAISan Francisco, CA
Enhances developer productivity for OpenAI's networking team by improving build systems, CI/CD pipelines, test harnesses, and workflows for C++ and Python codebases in multi-server environments. Requires experience with developer tools and infrastructure automation.