Software Engineer, GPU Infrastructure - HPC
Software engineer focused on reliability and uptime of OpenAI's GPU/HPC compute fleet through automation, monitoring tools, and performance optimization. Requires proficiency in Python/Go, Linux, networking, and data analysis skills.
About the job
In this role, you will:
- Build and maintain automation systems for provisioning and managing server fleets.
- Develop tools to monitor server health, performance, and lifecycle events.
- Collaborate with clusters, networking, and infrastructure teams.
- Partner with external operators to ensure a high level of quality.
- Identify and fix performance bottlenecks and inefficiencies.
- Continuously improve automation to reduce manual work.
You might thrive in this role if you have:
- Experience managing large-scale server environments.
- A balance of strengths in building and operationalizing.
- Proficiency in Python, Go, or similar languages.
- Strong Linux, networking, and server hardware knowledge.
- Comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool.
Prior hardware expertise is not required for this role.
Bonus Skills:
- Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning)
- Knowledge of hardware management protocols (e.g., IPMI, Redfish).
- High-performance computing (HPC) or distributed systems experience.
- Prior experience developing, managing, or designing hardware.
- Familiarity with monitoring tools (e.g., Prometheus, Grafana).
Skills
Python, Go, Linux, SQL, Promql, pandas, Prometheus, Grafana, Kubernetes, InfiniBand
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.