Research Engineer
Builds QA systems, tooling, and workflows to audit and validate large-scale RL training data from suppliers. Partners with vendors to improve data quality using Python, Docker, and AI/ML techniques for frontier AI infrastructure.
About the job
Responsibilities
- Define and enforce quality standards for training data
- Build tooling and workflows to audit supplier-generated datasets, including sampling strategies, validation pipelines (rule-based and model-assisted), and feedback loops
- Determine if and how human-in-the-loop review workflows can be used to optimize QA
- Partner with data vendors to debug quality issues, provide actionable feedback, and improve their data generation processes
- Continuously integrate QA learnings into infrastructure tools and data vendor portal to reduce anomalies, inconsistencies, and edge cases
Experience
- Proficiency in Python, Docker, and Linux environments
- Worked with large-scale datasets
- Evidence of rapid learning and adaptability in technical environments (e.g., programming competitions)
- Startup experience in early-stage technology companies with ability to work independently in fast-paced environments
- Familiarity with current AI tools and LLM capabilities
- Strong communication skills for remote collaboration across time zones
Strong candidates may also
- Understand common failure modes in training data
- Have experience building data validation pipelines and/or human-in-the-loop review systems
- Be detail-oriented and able to spot subtle inconsistencies or edge cases in data
- Be comfortable designing metrics, experiments, and QA processes, not just executing them
Skills
Python, Docker, Linux, LLMs, Data Validation, Large-Scale Datasets, Validation Pipelines, Human-In-The-Loop, AI Tools
Similar jobs
ML Engineering jobsResearch Engineer who productionizes robotics, perception, and machine-learning prototypes for reliable, real-time execution on hardware. Requires 3+ years of systems or production ML experience, strong modern C++, CUDA, Python, Linux, and performance optimization skills.
Senior software engineer building scalable backend and distributed systems for agentic AI products, integrating models into reliable customer-facing services. Requires 5+ years of production software engineering experience and strong backend, cloud, API, and AI systems expertise.
Build and operate LLM-powered agents that automate business workflows, integrating systems of record through MCP and validating behavior with structured evaluations. Requires 5+ years of software development experience, strong Python and SQL, and hands-on experience shipping agents to production.
Build and operate infrastructure-first production ML systems, including pipelines, model serving, CI/CD, monitoring, and automated retraining. The role requires 5+ years of production ML experience, strong Python and platform engineering skills, and hands-on expertise with cloud, Kubernetes, and MLOps tooling.
Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.