# Software Engineer, GPU Infrastructure - ChatGPT Engineering

**Company:** [OpenAI](https://hotfix.jobs/companies/openai)
**Location:** London, United Kingdom
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Go, Python, C++, Rust, Gpu Infrastructure, Distributed Systems, High-Performance Computing, ML Infrastructure, Cluster Orchestration, Observability, Monitoring, Capacity Planning, Incident Management, System Design, Production Engineering
**Posted:** 2026-08-12

> Build and operate software systems that manage the GPU fleet powering ChatGPT inference, including fleet health, capacity planning, resource utilization, and operational automation. The role requires 5+ years of production infrastructure experience and strong programming and distributed-systems skills.

## Job Description

## Responsibilities
- Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference.
- Develop systems for capacity planning, fleet health monitoring, and resource utilization.
- Automate operational workflows, including incident detection, diagnosis, and response.
- Identify and address bottlenecks affecting fleet reliability, scalability, and performance.
- Partner with infrastructure, research, and product engineering teams to improve the compute platform.

## Requirements
- Five or more years of software engineering experience building production infrastructure.
- Strong programming skills in Go, Python, C++, Rust, or a comparable language.
- Experience designing or operating highly available distributed systems.
- Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms.
- Strong debugging, systems design, and operational problem-solving skills.
- Strong communication skills and experience collaborating across engineering teams.
- Experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems.
- Background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering.
- Experience building software that automates operational workflows and reduces manual work.
- Experience with distributed infrastructure, cluster orchestration, or large-scale internal infrastructure platforms.
- Understanding of infrastructure observability, monitoring, capacity planning, and incident management.
- Ability to work across software engineering and systems operations, with ownership from design through production.
- Comfort working in fast-moving environments with significant technical ambiguity.

## Similar jobs

- [Software Engineer: Resiliency - Deploy at Scale](https://hotfix.jobs/jobs/21bb6f4f-daea-46ab-bac8-fff083d17451) - Cloudflare - London, United Kingdom
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Site Reliability Engineer, Infrastructure Platforms](https://hotfix.jobs/jobs/e28d2ba0-df16-41d6-9c60-e694ee353996) - GitLab - Remote
- [Member of Technical Staff](https://hotfix.jobs/jobs/d6c912e7-2f16-4a86-8738-980dd0b47cd0) - Perplexity - Remote - $220k – $405k/yr
- [Infrastructure Engineer](https://hotfix.jobs/jobs/e7b74501-38ae-4cd1-8b26-91e042b41460) - Writer - London, United Kingdom

**Apply:** https://hotfix.jobs/jobs/09f129a4-54df-4621-b46c-70bfd4bd992f
**Canonical:** https://hotfix.jobs/jobs/09f129a4-54df-4621-b46c-70bfd4bd992f