# Software Engineer, Pretraining

**Company:** [Cursor](https://hotfix.jobs/companies/cursor)
**Location:** San Francisco, CA
**Role:** Data Engineering
**Skills:** Distributed Systems, Data Pipelines, Data Platforms, Web Crawling, Data Quality, Machine Learning Models, Data Orchestration, Html Parsing, Multimodal Data, Observability
**Posted:** 2026-08-24

> Build the data systems behind frontier coding models, spanning web crawling, large-scale pipelines, data quality models, and training-ready dataset infrastructure. The role requires strong distributed-systems or data-platform expertise and end-to-end ownership.

## Job Description

## Responsibilities

### Data Quality
- Build and own high-throughput, fully telemetered data pipelines with end-to-end traceability.
- Train and ship high-throughput models for classifying, ranking, filtering, cleaning, and identifying data.
- Design and run scaling-ladder experiments on data mixtures, repeatability, and quality depth.
- Partner with data acquisition and training teams to identify missing or low-quality sources and measure their impact.
- Write performance-critical code and design rigorous experiments.

### Data Platform
- Build platforms that transform raw web, code, multimodal, and acquired data into training-ready datasets.
- Own pipelines, orchestration, and tooling for fast, reliable, observable, and reproducible pretraining data iteration.
- Create signals for data quality, lineage, freshness, and pipeline health.
- Partner with initial training, crawling, data quality, and acquisition teams to turn data ideas into measurable improvements.

### Crawling
- Build and scale web-crawling systems that discover, schedule, fetch, and parse documents across the open web.
- Improve URL seeding, scoring, and fair host scheduling.
- Increase crawl success and parsing quality by addressing antibot failures and improving extractors.
- Debug and harden crawl infrastructure for availability, recovery, and ingestion lag.
- Automate delivery of crawl datasets into downstream data pipelines.

## Requirements
- Strong infrastructure or data-platform background, ideally with significant depth in a specialized area.
- Demonstrated ability to architect and ship end-to-end systems with high ownership.
- Ability to debug complex systems independently and work alongside AI agents.
- Strong intuition for large-scale distributed systems.
- Interest in how pretraining data shapes model quality.

## Similar jobs

- [Analytics Engineer](https://hotfix.jobs/jobs/822824cc-1e77-4173-9f29-dc91ae0bd859) - Bevi - Boston, MA - $134k – $165k/yr
- [Data Engineer, GTM](https://hotfix.jobs/jobs/c0f575bd-fb5a-4a2b-aa6f-1e92c49eecc3) - Anthropic - San Francisco, CA - $320k – $405k/yr
- [Distributed Systems Engineer, Analytical Database Platform](https://hotfix.jobs/jobs/81ca3fba-abf2-4e8e-9e70-a484734b6cc9) - Cloudflare - Atlanta, GA - $54k – $91k/yr
- [Open Science Data Steward](https://hotfix.jobs/jobs/6513fd8c-ddf4-4a8b-b54d-ed97d14710b3) - Astera - Emeryville, CA - $100k – $150k/yr
- [Analytics Engineer, Data Platform](https://hotfix.jobs/jobs/e7f2755f-9557-483a-abd1-b4dae0399143) - Upside - Washington, DC - $149k – $180k/yr

**Apply:** https://hotfix.jobs/jobs/fa2181c4-874a-49b5-a05c-d1a4c4f29486
**Canonical:** https://hotfix.jobs/jobs/fa2181c4-874a-49b5-a05c-d1a4c4f29486