Software Engineer, Pretraining
Build the data systems behind frontier coding models, spanning web crawling, large-scale pipelines, data quality models, and training-ready dataset infrastructure. The role requires strong distributed-systems or data-platform expertise and end-to-end ownership.
About the job
Responsibilities
Data Quality
- Build and own high-throughput, fully telemetered data pipelines with end-to-end traceability.
- Train and ship high-throughput models for classifying, ranking, filtering, cleaning, and identifying data.
- Design and run scaling-ladder experiments on data mixtures, repeatability, and quality depth.
- Partner with data acquisition and training teams to identify missing or low-quality sources and measure their impact.
- Write performance-critical code and design rigorous experiments.
Data Platform
- Build platforms that transform raw web, code, multimodal, and acquired data into training-ready datasets.
- Own pipelines, orchestration, and tooling for fast, reliable, observable, and reproducible pretraining data iteration.
- Create signals for data quality, lineage, freshness, and pipeline health.
- Partner with initial training, crawling, data quality, and acquisition teams to turn data ideas into measurable improvements.
Crawling
- Build and scale web-crawling systems that discover, schedule, fetch, and parse documents across the open web.
- Improve URL seeding, scoring, and fair host scheduling.
- Increase crawl success and parsing quality by addressing antibot failures and improving extractors.
- Debug and harden crawl infrastructure for availability, recovery, and ingestion lag.
- Automate delivery of crawl datasets into downstream data pipelines.
Requirements
- Strong infrastructure or data-platform background, ideally with significant depth in a specialized area.
- Demonstrated ability to architect and ship end-to-end systems with high ownership.
- Ability to debug complex systems independently and work alongside AI agents.
- Strong intuition for large-scale distributed systems.
- Interest in how pretraining data shapes model quality.
Skills
Distributed Systems, Data Pipelines, Data Platforms, Web Crawling, Data Quality, Machine Learning Models, Data Orchestration, Html Parsing, Multimodal Data, Observability
Similar jobs
Data Engineering jobsBuild and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.
Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.