Skip to content
OpenAIOpenAI

Software Engineer, GPU Infrastructure - ChatGPT Engineering

Build and operate software systems that manage the GPU fleet powering ChatGPT inference, including fleet health, capacity planning, resource utilization, and operational automation. The role requires 5+ years of production infrastructure experience and strong programming and distributed-systems skills.

About the job

Responsibilities

  • Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference.
  • Develop systems for capacity planning, fleet health monitoring, and resource utilization.
  • Automate operational workflows, including incident detection, diagnosis, and response.
  • Identify and address bottlenecks affecting fleet reliability, scalability, and performance.
  • Partner with infrastructure, research, and product engineering teams to improve the compute platform.

Requirements

  • Five or more years of software engineering experience building production infrastructure.
  • Strong programming skills in Go, Python, C++, Rust, or a comparable language.
  • Experience designing or operating highly available distributed systems.
  • Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms.
  • Strong debugging, systems design, and operational problem-solving skills.
  • Strong communication skills and experience collaborating across engineering teams.
  • Experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems.
  • Background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering.
  • Experience building software that automates operational workflows and reduces manual work.
  • Experience with distributed infrastructure, cluster orchestration, or large-scale internal infrastructure platforms.
  • Understanding of infrastructure observability, monitoring, capacity planning, and incident management.
  • Ability to work across software engineering and systems operations, with ownership from design through production.
  • Comfort working in fast-moving environments with significant technical ambiguity.

Skills

Go, Python, C++, Rust, Gpu Infrastructure, Distributed Systems, High-Performance Computing, ML Infrastructure, Cluster Orchestration, Observability, Monitoring, Capacity Planning, Incident Management, System Design, Production Engineering

Cloudflare

Cloudflare

London, United Kingdom

Software Engineer: Resiliency - Deploy at Scale
No salary listedHybrid4+ YOEDevOps / SRE

Build and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

GitLab

GitLab

United Kingdom

Site Reliability Engineer, Infrastructure Platforms
No salary listedRemote5+ YOEDevOps / SRE

Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Writer

Writer

London, United Kingdom

Infrastructure Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.