Research Scientist - Post Training
Design and run SFT/RL experiments to measure dataset impact on LLM performance, capabilities, and alignment. Collaborate with labs to provide evidence of improvements; requires strong LLM training knowledge and fast experimentation, ideally undergrad/master's research background.
About the job
What You'll Do
- Run controlled SFT and RL experiments to measure the impact of our datasets on model performance.
- Quantify lift across capabilities (reasoning, tool use, long-horizon tasks, domain-specific workflows).
- Communicate your findings with partner labs to help drive sales.
- Work with internal SPLs to iterate on data quality based on your results.
What We're Looking For
- Strong familiarity with LLM training and evaluation methodologies.
- Genuine obsession with how data structure, selection, and quality drive model behavior.
- Ability to design lightweight experiments, move fast, and extract actionable insights from messy results.
- Comfort working across domains (you'll touch finance, software engineering, policy, and more).
- A bias toward building over theorizing.
- Great candidates are undergrad research or master's research (but haven't done a phd).
Compensation Structure
$250k-450k total compensation + equity
Skills
Llm Training, Sft, Rl, Model Evaluation, PyTorch, Hugging Face, Experiment Design, Data Quality, Reasoning Capabilities, Tool Use
Similar jobs
AI Research jobsConduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.
Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.
Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.
Research Engineer focused on designing benchmarks, evaluation systems, rubrics, and failure-analysis workflows for frontier language models. The role requires strong applied AI research and coding experience, with expertise in model evaluation, data quality, and backend systems.
Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.