Research Scientist, Interpretability
Conducts mechanistic interpretability research to reverse-engineer language models, developing methods to understand neural network algorithms for AI safety. Requires scientific research background, Python proficiency, and collaborative engineering mindset.
About the job
Responsibilities
- Develop methods for understanding LLMs by reverse engineering algorithms learned in their weights
- Design and run robust experiments, both quickly in toy scenarios and at scale in large models
- Create and analyze new interpretability features and circuits to better understand how models work
- Build infrastructure for running experiments and visualizing results
- Work with colleagues to communicate results internally and publicly
You may be a good fit if you
- Have a strong track record of scientific research (in any field), and have done some work on Interpretability
- Enjoy team science – working collaboratively to make big discoveries
- Are comfortable with messy experimental science. We're inventing the field as we work, and the first textbook is years away
- You view research and engineering as two sides of the same coin. Every team member writes code, designs and runs experiments, and interprets results
- You can clearly articulate and discuss the motivations behind your work, and teach us about what you've learned. You like writing up and communicating your results, even when they're null
Familiarity with Python is required.
Education requirements: At least a Bachelor's degree in a related field or equivalent experience.
Skills
Python, Mechanistic Interpretability, Neural Networks, LLMs, Transformer Circuits, Experiment Design, Reverse Engineering, Scientific Research, Data Visualization, Infrastructure Development
Similar jobs
AI Research jobsResearch Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.
Research role focused on improving agentic coding capabilities through reinforcement-learning training, synthetic data, coding environments, reward design, and evaluations. Requires strong Python engineering, scalable distributed-training experience, and a bachelor’s degree or equivalent; research experience and a PhD are preferred.
Conduct AI safety research across data curation, post-training, evaluations, synthetic data, and red-teaming to improve model reliability on harmful and dual-use requests. The role requires AI safety experience, Python, deep learning frameworks, and scalable technical research skills.
Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.
Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.