Research, Safety
Conduct AI safety research across data curation, post-training, evaluations, synthetic data, and red-teaming to improve model reliability on harmful and dual-use requests. The role requires AI safety experience, Python, deep learning frameworks, and scalable technical research skills.
About the job
Responsibilities
- Investigate how models handle harmful, sensitive, and dual-use requests, including what they learn from data and how training shapes reliable refusal and engagement boundaries.
- Build data-filtering pipelines and quality classifiers for pre-training corpora, and study downstream safety effects.
- Apply post-training methods, including RLHF, RLAIF, and policy-based reasoning approaches.
- Design, build, and maintain safety evaluations for long-horizon and agentic tasks.
- Generate and curate synthetic data for training and evaluating refusal boundaries and safety-relevant behaviors.
- Red-team models and products to identify failure modes, jailbreaks, and emergent risks, then design mitigations.
Requirements
- Bachelor's degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline.
- Background in AI safety research and hands-on experience with at least one of RLHF/RLAIF, alignment and preference modeling, deliberative alignment, safety evaluations, or red-teaming.
- Proficiency in Python and familiarity with deep learning frameworks such as PyTorch, TensorFlow, or JAX.
- Experience debugging distributed training and writing scalable code.
- Strong written communication and ability to explain complex technical concepts.
Nice-to-haves
- Experience evaluating long-horizon, multi-step, or agentic tasks.
- Experience generating synthetic data at scale.
- Experience with modern red-teaming and jailbreaking techniques.
- AI safety research contributions, including publications, open-source evaluations, or public red-teaming work.
- Familiarity with scalable oversight, reward hacking, and jailbreak robustness.
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline, or equivalent industry research experience.
Compensation & Benefits
- Annual salary range: $350,000–$475,000 USD.
- Health, dental, and vision benefits.
- Unlimited PTO.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship available.
Skills
Python, PyTorch, TensorFlow, JAX, RLHF, Rlaif, Preference Modeling, Safety Evaluations, Red-Teaming, Synthetic Data, Distributed Training, Jailbreaking, Scalable Oversight, Reward Hacking, Agentic Tasks
Similar jobs
AI Research jobsResearch Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.
Research role focused on improving agentic coding capabilities through reinforcement-learning training, synthetic data, coding environments, reward design, and evaluations. Requires strong Python engineering, scalable distributed-training experience, and a bachelor’s degree or equivalent; research experience and a PhD are preferred.
Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.
Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.
Conducts hands-on medicinal chemistry research to evaluate AI-generated molecules and synthetic routes, advancing small-molecule programs from design through experimental validation. Requires a chemistry PhD, sustained synthetic experience, and cross-functional collaboration skills.