Skip to content
Distyl AIDistyl AI

Research Engineer, Data

Research Engineers at Distyl build data systems and pipelines that power reliable compound AI workflows in enterprise environments. They create data quality frameworks, synthetic data strategies, and evaluation tools while partnering with researchers and customers to turn raw data into production AI value.

About the job

Key Responsibilities

  • Design and build data systems that power reliable AI workflows across enterprise environments
  • Develop pipelines for collecting, cleaning, transforming, labeling, and evaluating domain-specific data used by AI systems
  • Create data quality frameworks that identify coverage gaps, ambiguity, drift, duplication, leakage, and other failure modes
  • Build tools and workflows that help teams turn raw customer data into usable context for retrieval, evaluation, reasoning, and execution
  • Partner with AI Researchers and AI Engineers to understand how data quality affects system behavior and production outcomes
  • Develop synthetic data, annotation, and feedback-loop strategies to improve system performance in areas where real-world data is sparse or noisy
  • Analyze customer workflows and datasets to determine what information AI systems need, where that information should come from, and how it should be represented
  • Communicate clearly with internal teams and customer stakeholders about data assumptions, limitations, risks, and tradeoffs

Requirements

  • Experience building data pipelines, evaluation datasets, labeling workflows, retrieval corpora, or similar systems that improve model or agent behavior
  • Strong data engineering fundamentals: clean Python and SQL, data modeling, pipeline reliability, maintainable production systems
  • Research-oriented builder comfortable investigating how data quality, structure, and representation affect AI system performance
  • Use AI tools daily to accelerate coding, analysis, debugging, exploration, and workflow automation
  • Comfort reasoning through messy enterprise datasets, incomplete documentation, conflicting business definitions, and changing requirements
  • Bias towards measurement: make data quality and system behavior observable through concrete metrics, evaluations, and experiments
  • Ability to work directly with customer teams to understand their data, ask precise questions, and explain tradeoffs clearly
  • Ownership mentality for whether the data layer enables the AI system to deliver reliable value in production

Compensation and Benefits

  • Base salary range: $150,000 – $250,000 (depending on experience, location, and level)
  • Meaningful equity
  • 100% covered medical, dental, and vision for employees and dependents
  • 401(k) with additional perks (e.g., commuter benefits, in-office lunch)
  • Access to state-of-the-art models, generous usage of modern AI tools, and real-world business problems
  • Ownership of high-impact projects across top enterprises
  • Mission-driven, fast-moving culture that prizes curiosity, pragmatism, and excellence

Skills

Python, SQL, Data Pipelines, Data Quality Frameworks, Synthetic Data, Evaluation Datasets, Labeling Workflows, Retrieval Corpora, Ai Systems, Data Modeling

Upside

Upside

Washington, DC
Analytics Engineer, Data Platform
$149k+/yrHybrid3+ YOEData Engineering

Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.

Benchling

Benchling

San Francisco, CA

Data Engineer
$153k+/yrHybrid3+ YOEData Engineering

Build and operate reliable, production-grade data pipelines, warehouse infrastructure, and trusted datasets supporting company-wide analytics and AI initiatives. The role requires 3+ years of production data engineering experience, strong SQL and Python skills, and experience with Snowflake, dbt, cloud infrastructure, and orchestration.

Sigma

Sigma

New York, NY

Data Engineer
$140k+/yrOn-site3+ YOEData Engineering

Build and scale data architecture, governance, and ETL pipelines across Snowflake, Databricks, and cloud platforms. The role requires 3+ years of data engineering experience, strong API and pipeline expertise, and the ability to collaborate across technical and business teams.

OnePay

OnePay

United States

Analytics Engineer
$140k+/yrRemote3+ YOEData Engineering

The Analytics Engineer will own OnePay’s analytical data foundation, building trusted dbt models, tests, documentation, dashboards, and semantic metrics on Databricks. The role requires at least three years of analytics engineering experience, expert SQL and dbt skills, strong data-quality practices, and hands-on use of AI coding tools.

Stuut

Stuut

San Francisco, CA
Data Engineer
$135k+/yrOn-site3+ YOEData Engineering

Build and own Stuut’s foundational data platform, including ingestion pipelines, canonical models, semantic layers, and observability. The role requires 3+ years of production data pipeline experience with Python, SQL, cloud warehouses, and ETL/ELT tooling.