Skip to content

Research Engineer - Data Infrastructure

Build the data infrastructure, processing pipelines, curation strategies, and tooling that power frontier AI model training. The role requires strong distributed-data engineering experience and the ability to measure how data quality affects model outcomes.

About the job

Responsibilities

  • Build large-scale data pipelines for collecting, processing, filtering, and transforming datasets used to train state-of-the-art models.
  • Train models for data-processing pipelines, including classifiers, quality filters, and labeling models.
  • Design data-curation strategies such as deduplication, quality scoring, labeling, and augmentation to improve model performance.
  • Create tooling and infrastructure that enables researchers to explore and train on massive datasets quickly and reliably.

Requirements

  • Demonstrated ability to solve difficult engineering problems through past projects, designs, or GitHub contributions.
  • Experience building data-intensive systems, ideally supporting machine-learning training pipelines.
  • Strong engineering skills in distributed data processing at scale, including Kubernetes or custom pipelines over large datasets.
  • Ability to independently evaluate how data quality, composition, and curation affect model outcomes and build tooling to measure those effects.

Nice-to-haves

  • Experience building or operating web crawlers.

Compensation and Benefits

  • Annual professional-development stipend.
  • Annual stipend for social travel with colleagues.
  • Annual company offsite.
  • Monthly co-working stipend for employees outside main hubs.

Skills

Data Pipelines, Distributed Data Processing, Machine Learning, Kubernetes, Data Curation, Deduplication, Quality Scoring, Data Labeling, Data Augmentation, Web Crawlers

Bevi

Bevi

Boston, MA

Analytics Engineer
$134k+/yrHybrid4+ YOEData Engineering

Build and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.

Anthropic

Anthropic

San Francisco, CA
Data Engineer, GTM
$320k+/yrHybrid5+ YOEData Engineering

Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.

Cloudflare

Cloudflare

Atlanta, GA
Distributed Systems Engineer, Analytical Database Platform
$54k+/yrHybrid3+ YOEData Engineering

Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.

Astera

Astera

Emeryville, CA

Open Science Data Steward
$100k+/yrOn-site3+ YOEData Engineering

Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.

Upside

Upside

Washington, DC
Analytics Engineer, Data Platform
$149k+/yrHybrid3+ YOEData Engineering

Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.