Skip to content

AI Infrastructure Engineer

Operates and improves the infrastructure powering large-scale post-training and reinforcement learning runs, partnering with researchers to debug failures, improve reliability, and automate recovery. Requires 4+ years operating distributed production systems and strong Python, Go, or C++ skills.

About the job

Responsibilities

  • Own the reliability, performance, and uptime of large-scale post-training and reinforcement learning training jobs from launch through completion.
  • Partner with research teams during active model runs to unblock training and accelerate iteration.
  • Debug failures across accelerators, networking, storage, schedulers, and training frameworks, driving issues to root cause.
  • Build monitoring, alerting, and automated recovery systems so runs self-heal or fail fast.
  • Improve checkpointing, fault tolerance, and job scheduling to minimize compute lost to hardware failures.
  • Build internal tools that reduce toil and improve cluster utilization across post-training and RL workloads.
  • Participate in an on-call rotation supporting production model runs.
  • Write postmortems and convert recurring failure patterns into permanent infrastructure fixes.

Requirements

  • 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production.
  • Experience debugging complex distributed-systems failures involving networking, hardware, kernels, or schedulers.
  • Strong software engineering skills in Python and/or Go/C++, with sound judgment about when to script a fix versus build a system.
  • Strong foundation in Linux systems internals and networking fundamentals.
  • Comfort owning production systems and participating in on-call rotations.

Nice-to-haves

  • Experience operating GPU or TPU training clusters at scale.
  • Familiarity with post-training and reinforcement learning techniques, including RLHF, PPO, and DPO.
  • Experience with reward model serving, rollout generation, and mixed training/inference workloads.
  • Experience with distributed training frameworks such as PyTorch and Ray.
  • Experience with job schedulers such as Slurm and Kubernetes.
  • Experience with high-performance networking, including InfiniBand, RDMA, and NCCL.
  • Experience building observability tooling for ML training.
  • Experience working in fast-changing, research-driven environments.

Compensation

  • Expected annual salary: $350,000–$475,000 USD.
  • Health, dental, and vision benefits; unlimited PTO; paid parental leave; and relocation support.

Skills

Python, Go, C++, Linux, Networking, Gpu Clusters, Tpu, PyTorch, Ray, Slurm, Kubernetes, InfiniBand, Rdma, Nccl

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research, General Agents
$350k+/yrHybridML Engineering

Research-focused engineer advancing agentic model capabilities across synthetic data, task environments, evaluations, training, and usability improvements. Requires strong Python engineering, deep learning framework experience, scalable distributed training skills, and scientific experimentation ability.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research, RL Scaling
$350k+/yrHybridML Engineering

Researcher focused on scaling reinforcement learning for frontier models, with ownership spanning asynchronous RL algorithms, inference and distributed training systems, and large-scale empirical studies. Requires strong Python and deep learning experience, scalable systems debugging, and rigorous research judgment.

OpenAI

OpenAI

San Francisco, CA

Machine Learning Engineer, Multimodal Perception and Authentication
$342k+/yrHybridML Engineering

Develop multimodal perception and authentication systems combining visual, audio, and other sensor signals for real-world AI products. The role requires machine learning expertise, practical research-to-system experience, and proficiency in Python and PyTorch with comfort in C++.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Trainium
$295k+/yrHybrid3+ YOEML Engineering

Build and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.