# AI Infrastructure Engineer

**Company:** [Thinking Machines Lab](https://hotfix.jobs/companies/thinking-machines-lab)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $350k – $475k/yr
**Experience:** 4+ years
**Skills:** Python, Go, C++, Linux, Networking, Gpu Clusters, Tpu, PyTorch, Ray, Slurm, Kubernetes, InfiniBand, Rdma, Nccl
**Posted:** 2026-08-28

> Operates and improves the infrastructure powering large-scale post-training and reinforcement learning runs, partnering with researchers to debug failures, improve reliability, and automate recovery. Requires 4+ years operating distributed production systems and strong Python, Go, or C++ skills.

## Job Description

## Responsibilities
- Own the reliability, performance, and uptime of large-scale post-training and reinforcement learning training jobs from launch through completion.
- Partner with research teams during active model runs to unblock training and accelerate iteration.
- Debug failures across accelerators, networking, storage, schedulers, and training frameworks, driving issues to root cause.
- Build monitoring, alerting, and automated recovery systems so runs self-heal or fail fast.
- Improve checkpointing, fault tolerance, and job scheduling to minimize compute lost to hardware failures.
- Build internal tools that reduce toil and improve cluster utilization across post-training and RL workloads.
- Participate in an on-call rotation supporting production model runs.
- Write postmortems and convert recurring failure patterns into permanent infrastructure fixes.

## Requirements
- 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production.
- Experience debugging complex distributed-systems failures involving networking, hardware, kernels, or schedulers.
- Strong software engineering skills in Python and/or Go/C++, with sound judgment about when to script a fix versus build a system.
- Strong foundation in Linux systems internals and networking fundamentals.
- Comfort owning production systems and participating in on-call rotations.

## Nice-to-haves
- Experience operating GPU or TPU training clusters at scale.
- Familiarity with post-training and reinforcement learning techniques, including RLHF, PPO, and DPO.
- Experience with reward model serving, rollout generation, and mixed training/inference workloads.
- Experience with distributed training frameworks such as PyTorch and Ray.
- Experience with job schedulers such as Slurm and Kubernetes.
- Experience with high-performance networking, including InfiniBand, RDMA, and NCCL.
- Experience building observability tooling for ML training.
- Experience working in fast-changing, research-driven environments.

## Compensation
- Expected annual salary: **$350,000–$475,000 USD**.
- Health, dental, and vision benefits; unlimited PTO; paid parental leave; and relocation support.

## Similar jobs

- [Research Software Engineer, Post Training](https://hotfix.jobs/jobs/168e3c8f-8577-4482-bf94-91b3d11744ba) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Research, General Agents](https://hotfix.jobs/jobs/e075e229-9ae8-46ba-92a5-20a0a6d4f0db) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Research, RL Scaling](https://hotfix.jobs/jobs/bd1b548c-7c82-4627-84a9-149150dad06d) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Machine Learning Engineer, Multimodal Perception and Authentication](https://hotfix.jobs/jobs/26314e6e-423f-4a3f-b400-ee0b56429f64) - OpenAI - San Francisco, CA - $342k – $399k/yr
- [Software Engineer, Trainium](https://hotfix.jobs/jobs/5e1a8341-f0e9-45c8-8492-12782b38f079) - OpenAI - San Francisco, CA - $295k – $380k/yr

**Apply:** https://hotfix.jobs/jobs/0d5aa4cf-a861-427b-8d8f-cba8ac94104e
**Canonical:** https://hotfix.jobs/jobs/0d5aa4cf-a861-427b-8d8f-cba8ac94104e