Skip to content
hudhud

Research Engineer

Builds QA systems, tooling, and workflows to audit and validate large-scale RL training data from suppliers. Partners with vendors to improve data quality using Python, Docker, and AI/ML techniques for frontier AI infrastructure.

About the job

Responsibilities

  • Define and enforce quality standards for training data
  • Build tooling and workflows to audit supplier-generated datasets, including sampling strategies, validation pipelines (rule-based and model-assisted), and feedback loops
  • Determine if and how human-in-the-loop review workflows can be used to optimize QA
  • Partner with data vendors to debug quality issues, provide actionable feedback, and improve their data generation processes
  • Continuously integrate QA learnings into infrastructure tools and data vendor portal to reduce anomalies, inconsistencies, and edge cases

Experience

  • Proficiency in Python, Docker, and Linux environments
  • Worked with large-scale datasets
  • Evidence of rapid learning and adaptability in technical environments (e.g., programming competitions)
  • Startup experience in early-stage technology companies with ability to work independently in fast-paced environments
  • Familiarity with current AI tools and LLM capabilities
  • Strong communication skills for remote collaboration across time zones

Strong candidates may also

  • Understand common failure modes in training data
  • Have experience building data validation pipelines and/or human-in-the-loop review systems
  • Be detail-oriented and able to spot subtle inconsistencies or edge cases in data
  • Be comfortable designing metrics, experiments, and QA processes, not just executing them

Skills

Python, Docker, Linux, LLMs, Data Validation, Large-Scale Datasets, Validation Pipelines, Human-In-The-Loop, AI Tools

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

PathAI

PathAI

Boston, MA
Machine Learning Engineer III
$131k+/yrOn-site5+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.