Latest DevOps / SRE jobs at OpenAI
Job results
Build and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.
Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.
The Network Engineer will design, operate, troubleshoot, and automate secure enterprise and cloud networks across offices, labs, and production services. The role combines network architecture and lifecycle planning with incident response, operational delivery, and automation using APIs, Infrastructure as Code, Git, testing, and CI/CD.
Build and operate software systems that manage the GPU fleet powering ChatGPT inference, including fleet health, capacity planning, resource utilization, and operational automation. The role requires 5+ years of production infrastructure experience and strong programming and distributed-systems skills.
Build, scale, and operate OpenAI's global compute infrastructure for frontier AI models like GPT-5.6. Solve complex cross-disciplinary problems spanning distributed systems, hardware, ML infrastructure, power/cooling, manufacturing, supply chain, and data center development at unprecedented scale.
Build and operate scalable CI and Bazel-based build systems that accelerate engineering velocity and reliability for OpenAI's products and infrastructure.
Lead deployment and operations for OpenAI’s custom silicon and systems into data center environments. Drive hardware bring-up, validation, production deployment, and fleet reliability at scale while leading a technical team.
Builds infrastructure to monitor, detect, remediate, and verify hardware health across global GPU/CPU clusters at hyperscale. Owns node lifecycle workflows and partners with teams to ensure compute reliability for AI training and inference. Requires 7+ years experience with Python, distributed systems, and operational tooling.
Builds and improves CI/CD, testing, validation, and release tooling for OpenAI's inference runtime teams to ensure reliable, performant model deployments across ChatGPT, API, and research workloads. Requires strong Python skills, developer productivity experience, and high ownership in ambiguous environments.
Builds and operates high-performance networking infrastructure for OpenAI's large-scale AI training and inference, focusing on host networking, datacenter fabrics, and WAN systems. Optimizes latency, reliability, and scalability using technologies like RDMA, InfiniBand, and RoCE; requires strong systems programming in C++, Python, or Go.
Develops and maintains custom networking operating system firmware for AI supercomputers, integrating Linux kernel, switch ASICs, and control-plane services. Requires deep expertise in SONiC, SAI, routing protocols, and platform bring-up across hardware and software boundaries.
Optimizes performance across Codex AI system's stack including LLM inference, cloud orchestration, and agent behavior to reduce latency and costs. Collaborates with researchers and engineers on high-impact improvements in a high-ownership role.
Builds and improves developer tools, CI/CD pipelines, and testing workflows to boost productivity for OpenAI's model performance engineering teams. Requires strong Python skills, experience with developer infrastructure, and ability to work in ambiguous environments.
Enhances developer productivity for OpenAI's networking team by improving build systems, CI/CD pipelines, test harnesses, and workflows for C++ and Python codebases in multi-server environments. Requires experience with developer tools and infrastructure automation.
Builds systems and tooling to measure, monitor, and optimize token throughput from GPU infrastructure for OpenAI workloads. Integrates partner compute environments, benchmarks performance, analyzes tokenomics, and develops operational metrics and dashboards. Requires strong distributed systems and infrastructure engineering experience.
Builds and optimizes large-scale compute infrastructure for AI workloads, spanning hardware automation, distributed systems, Kubernetes orchestration, networking, storage, and developer tools. Requires strong systems engineering experience in performance, reliability, and production infrastructure.
Leads execution of CPU, storage, PoP, and WAN infrastructure programs to activate compute clusters and expand global networks. Requires 8+ years in technical program management with deep knowledge of hardware, networking, and data center deployments.
Designs, validates, and scales secure OT network architectures for high-density AI data centers, including controls systems, telemetry, and integration with IT infrastructure. Requires 8+ years in OT networking, industrial protocols, and resilient topologies in mission-critical environments.
Evaluates new hardware platforms by porting benchmarks and workloads, analyzes performance across compute/memory/networking, identifies bottlenecks, and optimizes for AI systems. Requires expertise in performance analysis, system architecture, and debugging across hardware/software boundaries.
Defines rack- and cluster-level reference architectures for AI infrastructure, translates workload requirements into designs, collaborates with partners and modeling teams to evaluate tradeoffs, and drives vendor roadmaps to address technology gaps.
Develop and maintain performance modeling tools to analyze AI system behavior, evaluate tradeoffs in compute, memory, networking, and storage. Requires 1-2 years experience in software engineering or systems analysis, strong programming, and analytical skills.
Develops and maintains performance modeling tools and frameworks to evaluate AI system behavior, analyze tradeoffs in compute, memory, networking, and storage. Collaborates with architects on simulations and insights for infrastructure design; requires strong software/modeling background and system architecture knowledge.
Develops kernel performance optimizations, AI-assisted tooling, and observability infrastructure for AI-native hardware. Requires strong low-level systems experience, kernel/accelerator expertise, and familiarity with AI workflows for engineering acceleration.
Performance Engineer optimizes infrastructure and application performance for ChatGPT and OpenAI API, focusing on latency, throughput, and efficiency at scale. Requires 7+ years in high-scale systems with expertise in profiling, tracing, and cross-layer optimizations.
Owns end-to-end production-critical infrastructure for analytics platform, building performant backend systems in Rust or C++ and operating distributed services at scale on Kubernetes. Requires strong systems experience in performance optimization, debugging, and on-call reliability.
Builds and scales infrastructure for OpenAI's experimentation platform, including low-latency configuration delivery, high-throughput data ingestion, and analytics systems handling billions of evaluations. Requires expertise in large-scale distributed systems, performance optimization, and operational excellence.
Builds and operates continuous deployment platforms for safe, rapid code rollouts across Kubernetes clusters and global regions. Focuses on progressive delivery, GitOps, automation, and AI-assisted workflows to boost developer productivity.
Operational owner for identity-connected SaaS platforms ensuring compliance, access governance, and reliable workflows. Builds automation, integrations, and controlled changes partnering with Security and Engineering teams. Requires SaaS/identity experience and compliance knowledge.
Build observability infrastructure and AI-powered tools for OpenAI's large-scale production systems, including logging, metrics, and debugging UIs. Requires experience with distributed systems, Kubernetes, AWS, and observability tools.
Software engineer focused on reliability and uptime of OpenAI's GPU/HPC compute fleet through automation, monitoring tools, and performance optimization. Requires proficiency in Python/Go, Linux, networking, and data analysis skills.
Designs and builds distributed failure detection, tracing, and observability systems for large-scale AI training jobs. Requires deep expertise in performance, distributed systems, hardware, networking, and low-level software engineering.
Designs and operates CI/CD pipelines and release infrastructure for multi-component systems including bootloaders, firmware, and OTA updates. Requires strong automation skills in Python/Bash, Linux expertise, and experience with build systems for embedded/consumer products.
Builds and maintains scalable, reliable infrastructure including testing tools, automation, and resource management platforms for AI systems. Collaborates cross-functionally to ensure high availability, performance, and fault tolerance in a fast-paced environment.
Builds and maintains cloud infrastructure abstractions for scalable, reliable product platforms like ChatGPT. Requires 5+ years in core infrastructure, Kubernetes at scale, and cloud abstractions; onsite in San Francisco with on-call duties.
Design, build, and operate a multi-tenant caching platform powering OpenAI's inference, identity, and products. Requires 5+ years in distributed systems with deep Redis/Memcached expertise and Kubernetes experience.
Builds and maintains foundational systems, tools, and processes to boost developer productivity and engineering velocity at OpenAI. Requires 5+ years engineering experience, including infrastructure tooling, with core tech like Kubernetes, Python, and Terraform. Onsite in SF HQ.
Designs, implements, and operates infrastructure systems for model training and deployment on a massive GPU fleet. Requires experience with hyperscale compute, Kubernetes, public clouds like Azure, and strong programming skills.
Builds and scales massive Kubernetes clusters for OpenAI's frontier supercomputers, automates bare-metal provisioning, and ensures reliability across data centers for AI model training. Requires expertise in distributed systems, Kubernetes operations, and infrastructure automation.