Skip to content

Latest DevOps / SRE jobs at OpenAI

38 jobs

Job results

OpenAI

OpenAI

San Francisco, CA

Systems Integration Engineer, Build Systems | Consumer Devices
$293k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.

OpenAI

OpenAI

San Francisco, CA

Network Engineer
$293k+/yrHybridDevOps / SRE

Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.

OpenAI

OpenAI

London, United Kingdom

Network Engineer
No salary listedHybridDevOps / SRE

The Network Engineer will design, operate, troubleshoot, and automate secure enterprise and cloud networks across offices, labs, and production services. The role combines network architecture and lifecycle planning with incident response, operational delivery, and automation using APIs, Infrastructure as Code, Git, testing, and CI/CD.

OpenAI

OpenAI

London, United Kingdom

Software Engineer, GPU Infrastructure - ChatGPT Engineering
No salary listedHybrid5+ YOEDevOps / SRE

Build and operate software systems that manage the GPU fleet powering ChatGPT inference, including fleet health, capacity planning, resource utilization, and operational automation. The role requires 5+ years of production infrastructure experience and strong programming and distributed-systems skills.

OpenAI

OpenAI

San Francisco, CA
Data Center Compute Infrastructure
$230k+/yrOn-site5+ YOEDevOps / SRE

Build, scale, and operate OpenAI's global compute infrastructure for frontier AI models like GPT-5.6. Solve complex cross-disciplinary problems spanning distributed systems, hardware, ML infrastructure, power/cooling, manufacturing, supply chain, and data center development at unprecedented scale.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Full-Stack — Developer Experience
$185k+/yrOn-site5+ YOEDevOps / SRE

Build and operate scalable CI and Bazel-based build systems that accelerate engineering velocity and reliability for OpenAI's products and infrastructure.

OpenAI

OpenAI

San Francisco, CA

Tech Lead, Deployment & Operations — Custom Infrastructure
$342k+/yrHybrid8+ YOEDevOps / SRE

Lead deployment and operations for OpenAI’s custom silicon and systems into data center environments. Drive hardware bring-up, validation, production deployment, and fleet reliability at scale while leading a technical team.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Frontier Systems
$250k+/yrOn-site7+ YOEDevOps / SRE

Builds infrastructure to monitor, detect, remediate, and verify hardware health across global GPU/CPU clusters at hyperscale. Owns node lifecycle workflows and partners with teams to ensure compute reliability for AI training and inference. Requires 7+ years experience with Python, distributed systems, and operational tooling.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Productivity - Inference Runtime
$230k+/yrOn-siteDevOps / SRE

Builds and improves CI/CD, testing, validation, and release tooling for OpenAI's inference runtime teams to ensure reliable, performant model deployments across ChatGPT, API, and research workloads. Requires strong Python skills, developer productivity experience, and high ownership in ambiguous environments.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Core Network Engineering
$230k+/yrOn-siteDevOps / SRE

Builds and operates high-performance networking infrastructure for OpenAI's large-scale AI training and inference, focusing on host networking, datacenter fabrics, and WAN systems. Optimizes latency, reliability, and scalability using technologies like RDMA, InfiniBand, and RoCE; requires strong systems programming in C++, Python, or Go.

OpenAI

OpenAI

San Francisco, CA

Networking Operating System Firmware Engineer
$266k+/yrHybridDevOps / SRE

Develops and maintains custom networking operating system firmware for AI supercomputers, integrating Linux kernel, switch ASICs, and control-plane services. Requires deep expertise in SONiC, SAI, routing protocols, and platform bring-up across hardware and software boundaries.

OpenAI

OpenAI

San Francisco, CA

Performance & Systems Engineer, Codex
$295k+/yrHybridDevOps / SRE

Optimizes performance across Codex AI system's stack including LLM inference, cloud orchestration, and agent behavior to reduce latency and costs. Collaborates with researchers and engineers on high-impact improvements in a high-ownership role.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Productivity - Model Performance
$230k+/yrOn-siteDevOps / SRE

Builds and improves developer tools, CI/CD pipelines, and testing workflows to boost productivity for OpenAI's model performance engineering teams. Requires strong Python skills, experience with developer infrastructure, and ability to work in ambiguous environments.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Productivity - Networking
$230k+/yrOn-siteDevOps / SRE

Enhances developer productivity for OpenAI's networking team by improving build systems, CI/CD pipelines, test harnesses, and workflows for C++ and Python codebases in multi-server environments. Requires experience with developer tools and infrastructure automation.

OpenAI

OpenAI

San Francisco, CA
Tokens-as-a-Service (Taas) Software Engineer
$293k+/yrHybridDevOps / SRE

Builds systems and tooling to measure, monitor, and optimize token throughput from GPU infrastructure for OpenAI workloads. Integrates partner compute environments, benchmarks performance, analyzes tokenomics, and develops operational metrics and dashboards. Requires strong distributed systems and infrastructure engineering experience.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Compute Infrastructure
$230k+/yrHybridDevOps / SRE

Builds and optimizes large-scale compute infrastructure for AI workloads, spanning hardware automation, distributed systems, Kubernetes orchestration, networking, storage, and developer tools. Requires strong systems engineering experience in performance, reliability, and production infrastructure.

OpenAI

OpenAI

San Francisco, CA
CPU/Storage/PoP-WAN Program Manager
$342k+/yrHybrid8+ YOEDevOps / SRE

Leads execution of CPU, storage, PoP, and WAN infrastructure programs to activate compute clusters and expand global networks. Requires 8+ years in technical program management with deep knowledge of hardware, networking, and data center deployments.

OpenAI

OpenAI

San Francisco, CA

Data Center Controls Network Engineer
$257k+/yrHybrid8+ YOEDevOps / SRE

Designs, validates, and scales secure OT network architectures for high-density AI data centers, including controls systems, telemetry, and integration with IT infrastructure. Requires 8+ years in OT networking, industrial protocols, and resilient topologies in mission-critical environments.

OpenAI

OpenAI

San Francisco, CA
Workload Porting & Performance Engineer
$342k+/yrHybridDevOps / SRE

Evaluates new hardware platforms by porting benchmarks and workloads, analyzes performance across compute/memory/networking, identifies bottlenecks, and optimizes for AI systems. Requires expertise in performance analysis, system architecture, and debugging across hardware/software boundaries.

OpenAI

OpenAI

San Francisco, CA
3P Architect
$342k+/yrHybridDevOps / SRE

Defines rack- and cluster-level reference architectures for AI infrastructure, translates workload requirements into designs, collaborates with partners and modeling teams to evaluate tradeoffs, and drives vendor roadmaps to address technology gaps.

OpenAI

OpenAI

San Francisco, CA
Performance Modeling Engineer ~2
$266k+/yrHybrid1+ YOEDevOps / SRE

Develop and maintain performance modeling tools to analyze AI system behavior, evaluate tradeoffs in compute, memory, networking, and storage. Requires 1-2 years experience in software engineering or systems analysis, strong programming, and analytical skills.

OpenAI

OpenAI

San Francisco, CA
Performance Modeling Engineer
$266k+/yrHybridDevOps / SRE

Develops and maintains performance modeling tools and frameworks to evaluate AI system behavior, analyze tradeoffs in compute, memory, networking, and storage. Collaborates with architects on simulations and insights for infrastructure design; requires strong software/modeling background and system architecture knowledge.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Kernel Performance & AI Tooling
$266k+/yrHybridDevOps / SRE

Develops kernel performance optimizations, AI-assisted tooling, and observability infrastructure for AI-native hardware. Requires strong low-level systems experience, kernel/accelerator expertise, and familiarity with AI workflows for engineering acceleration.

OpenAI

OpenAI

San Francisco, CA
ChatGPT Performance Engineer
$325k+/yrRemote7+ YOEDevOps / SRE

Performance Engineer optimizes infrastructure and application performance for ChatGPT and OpenAI API, focusing on latency, throughput, and efficiency at scale. Requires 7+ years in high-scale systems with expertise in profiling, tracing, and cross-layer optimizations.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Infrastructure - Analytics Platform
$230k+/yrHybridDevOps / SRE

Owns end-to-end production-critical infrastructure for analytics platform, building performant backend systems in Rust or C++ and operating distributed services at scale on Kubernetes. Requires strong systems experience in performance optimization, debugging, and on-call reliability.

OpenAI

OpenAI

Seattle, WA

Senior Software Engineer, Infrastructure
$293k+/yrHybridDevOps / SRE

Builds and scales infrastructure for OpenAI's experimentation platform, including low-latency configuration delivery, high-throughput data ingestion, and analytics systems handling billions of evaluations. Requires expertise in large-scale distributed systems, performance optimization, and operational excellence.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Delivery / CD
$210k+/yrOn-siteDevOps / SRE

Builds and operates continuous deployment platforms for safe, rapid code rollouts across Kubernetes clusters and global regions. Focuses on progressive delivery, GitOps, automation, and AI-assisted workflows to boost developer productivity.

OpenAI

OpenAI

San Francisco, CA

IT Solutions Engineer
$251k+/yrHybridDevOps / SRE

Operational owner for identity-connected SaaS platforms ensuring compliance, access governance, and reliable workflows. Builds automation, integrations, and controlled changes partnering with Security and Engineering teams. Requires SaaS/identity experience and compliance knowledge.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Observability
$255k+/yrOn-siteDevOps / SRE

Build observability infrastructure and AI-powered tools for OpenAI's large-scale production systems, including logging, metrics, and debugging UIs. Requires experience with distributed systems, Kubernetes, AWS, and observability tools.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, GPU Infrastructure - HPC
$230k+/yrOn-siteDevOps / SRE

Software engineer focused on reliability and uptime of OpenAI's GPU/HPC compute fleet through automation, monitoring tools, and performance optimization. Requires proficiency in Python/Go, Linux, networking, and data analysis skills.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Platform Systems
$310k+/yrOn-siteDevOps / SRE

Designs and builds distributed failure detection, tracing, and observability systems for large-scale AI training jobs. Requires deep expertise in performance, distributed systems, hardware, networking, and low-level software engineering.

OpenAI

OpenAI

San Francisco, CA

Release Engineer, Consumer Products
$293k+/yrHybridDevOps / SRE

Designs and operates CI/CD pipelines and release infrastructure for multi-component systems including bootloaders, firmware, and OTA updates. Requires strong automation skills in Python/Bash, Linux expertise, and experience with build systems for embedded/consumer products.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Reliability
$230k+/yrOn-siteDevOps / SRE

Builds and maintains scalable, reliable infrastructure including testing tools, automation, and resource management platforms for AI systems. Collaborates cross-functionally to ensure high availability, performance, and fault tolerance in a fast-paced environment.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Cloud Infrastructure
$230k+/yrOn-site5+ YOEDevOps / SRE

Builds and maintains cloud infrastructure abstractions for scalable, reliable product platforms like ChatGPT. Requires 5+ years in core infrastructure, Kubernetes at scale, and cloud abstractions; onsite in San Francisco with on-call duties.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Caching Infrastructure
$230k+/yrOn-site5+ YOEDevOps / SRE

Design, build, and operate a multi-tenant caching platform powering OpenAI's inference, identity, and products. Requires 5+ years in distributed systems with deep Redis/Memcached expertise and Kubernetes experience.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Developer Productivity
$210k+/yrOn-site5+ YOEDevOps / SRE

Builds and maintains foundational systems, tools, and processes to boost developer productivity and engineering velocity at OpenAI. Requires 5+ years engineering experience, including infrastructure tooling, with core tech like Kubernetes, Python, and Terraform. Onsite in SF HQ.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Fleet Infrastructure
$230k+/yrHybridDevOps / SRE

Designs, implements, and operates infrastructure systems for model training and deployment on a massive GPU fleet. Requires experience with hyperscale compute, Kubernetes, public clouds like Azure, and strong programming skills.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Frontier Clusters Infrastructure
$230k+/yrOn-siteDevOps / SRE

Builds and scales massive Kubernetes clusters for OpenAI's frontier supercomputers, automates bare-metal provisioning, and ensures reliability across data centers for AI model training. Requires expertise in distributed systems, Kubernetes operations, and infrastructure automation.