Research Scientist, Safety Post Training
Develop and apply post-training methods and interpretability techniques to improve safety and understanding of frontier AI systems. Requires 3+ years of ML experience, expertise in RL techniques like RLHF and DPO, and published research in generative AI.
About the job
Responsibilities
- Design and run post-training pipelines to study how training choices affect model safety, robustness, and alignment properties.
- Develop interpretability-informed evaluations that reveal how and why models produce unsafe, deceptive, or otherwise undesirable behaviors, and use those insights to guide targeted mitigations.
- Collaborate with policymakers, engineers, and other researchers to translate post-training and interpretability findings into actionable safety standards, evaluation benchmarks, and best practices.
Requirements
- Commitment to promoting safe, secure, and trustworthy AI deployments.
- Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches.
- A track record of published research in machine learning, particularly in generative AI.
- At least three years of experience addressing sophisticated ML problems, whether in a research setting or in product development.
- Strong written and verbal communication skills.
Nice to Haves
- Experience with mechanistic interpretability, probing, or other techniques for understanding model internals.
- Familiarity with red-teaming or adversarial evaluation of post-trained models.
- Experience studying failure modes introduced or masked by post-training, such as reward hacking, sycophancy, or alignment faking.
Compensation and Benefits
Compensation packages include base salary, equity, and benefits. Eligible roles receive comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. This role may be eligible for a commuter stipend.
Skills
RLHF, Dpo, Grpo, Post-Training, Mechanistic Interpretability, Red-Teaming, Machine Learning, Generative AI, Model Alignment, Ml Prototyping
Similar jobs
AI Research jobsBuild and ship agentic AI product experiences, internal automation, and customer-facing features across the stack. The role requires 5+ years of software engineering experience, hands-on experience with AI or LLM-powered products, Python proficiency, and strong autonomy.
Leads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.
Research Scientist defining and executing research on reliable long-horizon agents in enterprise environments. The role focuses on post-training and reinforcement learning, agent memory, evaluation, verification, and structured representations, combining hands-on experimentation with product delivery and publication.
Conduct rigorous people research and applied data science to evaluate talent programs, organizational health, and employee experiences. The role requires advanced expertise in research design, experimentation, measurement, causal inference, statistical modeling, and responsible handling of sensitive employee data.
Conducts causal inference research for financial market prediction and portfolio optimization, developing and validating models from research through live trading. Requires Ph.D.-level coursework, strong causal inference and statistics expertise, mathematical ability, and production Python skills.