# Infrastructure Engineer, Pre-training

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** San Francisco, CA
**Role:** Data Engineering
**Salary:** $500k – $850k/yr
**Experience:** 7+ years
**Skills:** Spark, Python, Rust, Distributed Systems, Data Pipelines, MLOps, Machine Learning Infrastructure, Tokenization, Deduplication, Data Quality Assurance, Fault-Tolerant Systems
**Posted:** 2026-09-04

> Build scalable, fault-tolerant data infrastructure and pipelines that transform web-scale corpora into training datasets for large language models. The role requires substantial distributed-systems experience, Apache Spark expertise, and strong Python or Rust skills.

## Job Description

## Responsibilities
- Design and implement highly performant, reproducible, and traceable data-processing infrastructure for large language model training.
- Develop and maintain scalable processing primitives such as tokenization, deduplication, and chunking.
- Build robust systems for data-quality assurance and validation at scale.
- Collaborate with research teams to implement novel data-processing architectures.
- Build and operate end-to-end data pipelines that transform raw web-scale corpora into training-ready datasets.
- Design distributed-computing architectures for web-scale data processing.
- Build scalable infrastructure for model-training data preparation.
- Develop fault-tolerant distributed-processing systems.
- Implement infrastructure components based on research requirements.

## Requirements
- At least 5 years of professional experience outside internships.
- Strong software engineering skills and experience building high-throughput, fault-tolerant distributed systems.
- Hands-on experience with distributed-computing frameworks, particularly Apache Spark.
- Excellent problem-solving skills and attention to detail.
- Strong communication and collaboration skills.
- Advanced degree in Computer Science or a related field.
- Experience with language-model training infrastructure.
- Background in data infrastructure, MLOps, or machine-learning infrastructure.

## Nice to Have
- Significant experience building high-throughput, fault-tolerant distributed systems.
- Expertise with Python and Rust.
- Passion for system reliability and performance.
- Comfort working with ambiguous requirements and evolving specifications.
- Ability to take ownership of problems and drive solutions independently.
- Interest in machine-learning research and its infrastructure requirements.
- Ability to balance technical excellence with practical delivery.

## Compensation
- Annual salary: $500,000–$850,000 USD.

## Similar jobs

- [Lead Data Platform Engineer - Enterprise, Data & AI](https://hotfix.jobs/jobs/208fa17a-1c41-4472-8810-13a7f9492d1f) - Zoox - Foster City, CA - $230k – $277k/yr
- [Senior Data Engineer](https://hotfix.jobs/jobs/8fa75910-7872-42b1-b79f-9939f03c2318) - Garner Health - New York, NY - $220k – $245k/yr
- [Senior Database Reliability Engineer](https://hotfix.jobs/jobs/cc2c14f6-9532-42d5-89b2-55071c1e06c4) - Prompt Health - Remote - $220k – $240k/yr
- [Analytics Engineering Manager](https://hotfix.jobs/jobs/389b68c7-35ae-41d3-96c4-5c9e848ab653) - OnePay - Remote - $220k – $260k/yr
- [Senior Software Engineer, Data Infrastructure](https://hotfix.jobs/jobs/2c3e5efe-6f60-4a27-b7ec-1795674ea20d) - Decagon - San Francisco, CA - $200k – $400k/yr

**Apply:** https://hotfix.jobs/jobs/447aa589-2e6c-4331-b4e2-7afdfd8aeafd
**Canonical:** https://hotfix.jobs/jobs/447aa589-2e6c-4331-b4e2-7afdfd8aeafd