Infrastructure Engineer, Pre-training
Build scalable, fault-tolerant data infrastructure and pipelines that transform web-scale corpora into training datasets for large language models. The role requires substantial distributed-systems experience, Apache Spark expertise, and strong Python or Rust skills.
About the job
Responsibilities
- Design and implement highly performant, reproducible, and traceable data-processing infrastructure for large language model training.
- Develop and maintain scalable processing primitives such as tokenization, deduplication, and chunking.
- Build robust systems for data-quality assurance and validation at scale.
- Collaborate with research teams to implement novel data-processing architectures.
- Build and operate end-to-end data pipelines that transform raw web-scale corpora into training-ready datasets.
- Design distributed-computing architectures for web-scale data processing.
- Build scalable infrastructure for model-training data preparation.
- Develop fault-tolerant distributed-processing systems.
- Implement infrastructure components based on research requirements.
Requirements
- At least 5 years of professional experience outside internships.
- Strong software engineering skills and experience building high-throughput, fault-tolerant distributed systems.
- Hands-on experience with distributed-computing frameworks, particularly Apache Spark.
- Excellent problem-solving skills and attention to detail.
- Strong communication and collaboration skills.
- Advanced degree in Computer Science or a related field.
- Experience with language-model training infrastructure.
- Background in data infrastructure, MLOps, or machine-learning infrastructure.
Nice to Have
- Significant experience building high-throughput, fault-tolerant distributed systems.
- Expertise with Python and Rust.
- Passion for system reliability and performance.
- Comfort working with ambiguous requirements and evolving specifications.
- Ability to take ownership of problems and drive solutions independently.
- Interest in machine-learning research and its infrastructure requirements.
- Ability to balance technical excellence with practical delivery.
Compensation
- Annual salary: $500,000–$850,000 USD.
Skills
Spark, Python, Rust, Distributed Systems, Data Pipelines, MLOps, Machine Learning Infrastructure, Tokenization, Deduplication, Data Quality Assurance, Fault-Tolerant Systems
Similar jobs
Data Engineering jobsLeads the architecture, scaling, security, governance, and cost optimization of enterprise and AI data platforms. Requires 10+ years of data or software engineering experience, with expertise in production data foundations, CI/CD, governance, security, and performance optimization.
Build and scale data pipelines, reusable datasets, and validation frameworks supporting business intelligence, marketing, and data science. The role requires strong Python and SQL skills, modern data-stack experience, and at least four years of software or data engineering experience.
Own the reliability, performance, observability, scalability, and cost efficiency of large Aurora MySQL production environments supporting healthcare applications. The role requires 6+ years of database engineering experience, deep MySQL and AWS expertise, and strong skills in automation, incident response, and query optimization.
Leads an analytics engineering team that transforms raw data into reliable, actionable insights for product, marketing, and operations. The role requires 7+ years in data or analytics engineering, management experience, and advanced SQL, Databricks, and dbt expertise.
Senior Data Infrastructure Engineer responsible for building and operating reliable, low-latency streaming and batch data systems that support AI products. Requires 5+ years of production data infrastructure experience and expertise with technologies such as Kafka, Flink, ClickHouse, and Terraform.