# ML Infrastructure Engineer

**Company:** [xAI](https://hotfix.jobs/companies/xai)
**Location:** Palo Alto, CA
**Role:** ML Engineering
**Salary:** $180k – $440k/yr
**Experience:** 7+ years
**Skills:** Python, C++, Rust, JAX, PyTorch, CUDA, nvidia drivers, slurm, Linux, Ansible, puppet
**Posted:** 2026-07-22

> Build and scale GPU compute infrastructure, training frameworks, data pipelines, and ML platforms to power large-scale training and inference at xAI. Requires strong distributed systems and ML infrastructure experience plus Python proficiency.

## Job Description

## Responsibilities
- Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses
- Developing data pipelines and integrating large-scale data, training, and inference systems
- Collaborating with ML teams to productionize models and ensure seamless integration across the stack
- Ensuring scalability, reliability, and efficiency of large-scale machine learning systems
- Working across the full stack to solve complex problems independently
- Mentoring junior engineers and contributing to the growth of the team

## Basic Qualifications
- Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or other quantitative discipline; or equivalent work experience
- 2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications
- 2+ years experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists
- Strong proficiency with Python and experience with compiled languages such as C++ or Rust

## Preferred Skills and Experience
- Deep familiarity with modern ML frameworks such as JAX or PyTorch
- Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking
- Comfortable with Linux systems and orchestration tools
- Experience with job schedulers (e.g., Slurm), configuration management (Puppet/Ansible), or related infrastructure tooling

## Compensation and Benefits
- $180,000 - $440,000 USD total compensation (base salary is just one part; package also includes equity, comprehensive medical, vision, and dental coverage, 401(k), disability insurance, life insurance, and perks)

## Similar roles

- [Senior Software Engineer](https://hotfix.jobs/jobs/f3c2480b-2287-4246-9415-4d3a727e5994) - Deepgram - Remote - $180k – $240k/yr
- [Software Engineer, ML Data](https://hotfix.jobs/jobs/c04a70cf-91a2-4b7d-9504-b41b76234939) - Liftoff - San Francisco, CA - $180k – $230k/yr
- [Senior Software Engineer, AI](https://hotfix.jobs/jobs/b5ff4a30-137e-4cb9-93aa-59afeae0fa12) - Ironclad - San Francisco, CA - $180k – $220k/yr
- [Senior Machine Learning Engineer](https://hotfix.jobs/jobs/4713b492-d636-407a-bb9b-2a7357fb074a) - Teleskope - New York, NY - $180k – $210k/yr
- [Senior Machine Learning Engineer](https://hotfix.jobs/jobs/cc221086-96c9-44b8-9848-3fc8a12e202b) - Arcade - San Francisco, CA - $180k – $300k/yr

**Apply:** https://hotfix.jobs/jobs/60a4dff1-7af4-474a-8a5b-b9251cf4f4b8
**Canonical:** https://hotfix.jobs/jobs/60a4dff1-7af4-474a-8a5b-b9251cf4f4b8