Software Engineer, GPU Infrastructure - ChatGPT Engineering
Build and operate software systems that manage the GPU fleet powering ChatGPT inference, including fleet health, capacity planning, resource utilization, and operational automation. The role requires 5+ years of production infrastructure experience and strong programming and distributed-systems skills.
About the job
Responsibilities
- Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference.
- Develop systems for capacity planning, fleet health monitoring, and resource utilization.
- Automate operational workflows, including incident detection, diagnosis, and response.
- Identify and address bottlenecks affecting fleet reliability, scalability, and performance.
- Partner with infrastructure, research, and product engineering teams to improve the compute platform.
Requirements
- Five or more years of software engineering experience building production infrastructure.
- Strong programming skills in Go, Python, C++, Rust, or a comparable language.
- Experience designing or operating highly available distributed systems.
- Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms.
- Strong debugging, systems design, and operational problem-solving skills.
- Strong communication skills and experience collaborating across engineering teams.
- Experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems.
- Background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering.
- Experience building software that automates operational workflows and reduces manual work.
- Experience with distributed infrastructure, cluster orchestration, or large-scale internal infrastructure platforms.
- Understanding of infrastructure observability, monitoring, capacity planning, and incident management.
- Ability to work across software engineering and systems operations, with ownership from design through production.
- Comfort working in fast-moving environments with significant technical ambiguity.
Skills
Go, Python, C++, Rust, Gpu Infrastructure, Distributed Systems, High-Performance Computing, ML Infrastructure, Cluster Orchestration, Observability, Monitoring, Capacity Planning, Incident Management, System Design, Production Engineering
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.