Skip to content
CursorCursor

Software Engineer, Pretraining

Build the data systems behind frontier coding models, spanning web crawling, large-scale pipelines, data quality models, and training-ready dataset infrastructure. The role requires strong distributed-systems or data-platform expertise and end-to-end ownership.

About the job

Responsibilities

Data Quality

  • Build and own high-throughput, fully telemetered data pipelines with end-to-end traceability.
  • Train and ship high-throughput models for classifying, ranking, filtering, cleaning, and identifying data.
  • Design and run scaling-ladder experiments on data mixtures, repeatability, and quality depth.
  • Partner with data acquisition and training teams to identify missing or low-quality sources and measure their impact.
  • Write performance-critical code and design rigorous experiments.

Data Platform

  • Build platforms that transform raw web, code, multimodal, and acquired data into training-ready datasets.
  • Own pipelines, orchestration, and tooling for fast, reliable, observable, and reproducible pretraining data iteration.
  • Create signals for data quality, lineage, freshness, and pipeline health.
  • Partner with initial training, crawling, data quality, and acquisition teams to turn data ideas into measurable improvements.

Crawling

  • Build and scale web-crawling systems that discover, schedule, fetch, and parse documents across the open web.
  • Improve URL seeding, scoring, and fair host scheduling.
  • Increase crawl success and parsing quality by addressing antibot failures and improving extractors.
  • Debug and harden crawl infrastructure for availability, recovery, and ingestion lag.
  • Automate delivery of crawl datasets into downstream data pipelines.

Requirements

  • Strong infrastructure or data-platform background, ideally with significant depth in a specialized area.
  • Demonstrated ability to architect and ship end-to-end systems with high ownership.
  • Ability to debug complex systems independently and work alongside AI agents.
  • Strong intuition for large-scale distributed systems.
  • Interest in how pretraining data shapes model quality.

Skills

Distributed Systems, Data Pipelines, Data Platforms, Web Crawling, Data Quality, Machine Learning Models, Data Orchestration, Html Parsing, Multimodal Data, Observability

Bevi

Bevi

Boston, MA

Analytics Engineer
$134k+/yrHybrid4+ YOEData Engineering

Build and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.

Anthropic

Anthropic

San Francisco, CA
Data Engineer, GTM
$320k+/yrHybrid5+ YOEData Engineering

Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.

Cloudflare

Cloudflare

Atlanta, GA
Distributed Systems Engineer, Analytical Database Platform
$54k+/yrHybrid3+ YOEData Engineering

Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.

Astera

Astera

Emeryville, CA

Open Science Data Steward
$100k+/yrOn-site3+ YOEData Engineering

Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.

Upside

Upside

Washington, DC
Analytics Engineer, Data Platform
$149k+/yrHybrid3+ YOEData Engineering

Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.