Research Engineer, Data
Research Engineers at Distyl build data systems and pipelines that power reliable compound AI workflows in enterprise environments. They create data quality frameworks, synthetic data strategies, and evaluation tools while partnering with researchers and customers to turn raw data into production AI value.
About the job
Key Responsibilities
- Design and build data systems that power reliable AI workflows across enterprise environments
- Develop pipelines for collecting, cleaning, transforming, labeling, and evaluating domain-specific data used by AI systems
- Create data quality frameworks that identify coverage gaps, ambiguity, drift, duplication, leakage, and other failure modes
- Build tools and workflows that help teams turn raw customer data into usable context for retrieval, evaluation, reasoning, and execution
- Partner with AI Researchers and AI Engineers to understand how data quality affects system behavior and production outcomes
- Develop synthetic data, annotation, and feedback-loop strategies to improve system performance in areas where real-world data is sparse or noisy
- Analyze customer workflows and datasets to determine what information AI systems need, where that information should come from, and how it should be represented
- Communicate clearly with internal teams and customer stakeholders about data assumptions, limitations, risks, and tradeoffs
Requirements
- Experience building data pipelines, evaluation datasets, labeling workflows, retrieval corpora, or similar systems that improve model or agent behavior
- Strong data engineering fundamentals: clean Python and SQL, data modeling, pipeline reliability, maintainable production systems
- Research-oriented builder comfortable investigating how data quality, structure, and representation affect AI system performance
- Use AI tools daily to accelerate coding, analysis, debugging, exploration, and workflow automation
- Comfort reasoning through messy enterprise datasets, incomplete documentation, conflicting business definitions, and changing requirements
- Bias towards measurement: make data quality and system behavior observable through concrete metrics, evaluations, and experiments
- Ability to work directly with customer teams to understand their data, ask precise questions, and explain tradeoffs clearly
- Ownership mentality for whether the data layer enables the AI system to deliver reliable value in production
Compensation and Benefits
- Base salary range: $150,000 – $250,000 (depending on experience, location, and level)
- Meaningful equity
- 100% covered medical, dental, and vision for employees and dependents
- 401(k) with additional perks (e.g., commuter benefits, in-office lunch)
- Access to state-of-the-art models, generous usage of modern AI tools, and real-world business problems
- Ownership of high-impact projects across top enterprises
- Mission-driven, fast-moving culture that prizes curiosity, pragmatism, and excellence
Skills
Python, SQL, Data Pipelines, Data Quality Frameworks, Synthetic Data, Evaluation Datasets, Labeling Workflows, Retrieval Corpora, Ai Systems, Data Modeling
Similar jobs
Data Engineering jobsOwn and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.
Build and operate reliable, production-grade data pipelines, warehouse infrastructure, and trusted datasets supporting company-wide analytics and AI initiatives. The role requires 3+ years of production data engineering experience, strong SQL and Python skills, and experience with Snowflake, dbt, cloud infrastructure, and orchestration.
Build and scale data architecture, governance, and ETL pipelines across Snowflake, Databricks, and cloud platforms. The role requires 3+ years of data engineering experience, strong API and pipeline expertise, and the ability to collaborate across technical and business teams.
The Analytics Engineer will own OnePay’s analytical data foundation, building trusted dbt models, tests, documentation, dashboards, and semantic metrics on Databricks. The role requires at least three years of analytics engineering experience, expert SQL and dbt skills, strong data-quality practices, and hands-on use of AI coding tools.
Build and own Stuut’s foundational data platform, including ingestion pipelines, canonical models, semantic layers, and observability. The role requires 3+ years of production data pipeline experience with Python, SQL, cloud warehouses, and ETL/ELT tooling.