Research Engineer, Machine Learning (RL Velocity)
Builds and optimizes RL training infrastructure, removes bottlenecks in the RL stack, and partners with researchers to accelerate model development at scale. Requires strong software engineering, ML infra experience, and comfort across the stack.
About the job
Responsibilities
- Build and improve the RL training infrastructure that researchers depend on day-to-day
- Identify and remove bottlenecks across the RL stack: debugging, profiling, and rearchitecting where needed
- Partner closely with researchers and with adjacent engineering teams (inference, sandboxing, and many more) to understand pain points and ship tooling that makes them faster
- Own the reliability and performance of research runs end-to-end
- Contribute to design decisions that shape how Anthropic does RL at scale
You may be a good fit if you
- Have strong software engineering fundamentals and a track record of building performant, reliable systems
- Have worked on ML infrastructure, distributed systems, or research tooling
- Care about enabling other people's work and find leverage through platforms rather than individual experiments
- Are comfortable operating across the stack, from low-level performance work to RL algorithms
- Have a bias toward shipping and iterating quickly, with a mix of high agency and low ego
Strong candidates may also have
- Experience with large-scale distributed training (RL, pre-training, or post-training)
- Familiarity with JAX, PyTorch, or similar ML frameworks
- A track record of operating at the edge of research and infra in a fast-moving environment
Logistics
Annual Salary: $500,000—$850,000 USD
Skills
Reinforcement Learning, JAX, PyTorch, Distributed Systems, ML Infrastructure, Research Tooling, Distributed Training
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Operates and improves the infrastructure powering large-scale post-training and reinforcement learning runs, partnering with researchers to debug failures, improve reliability, and automate recovery. Requires 4+ years operating distributed production systems and strong Python, Go, or C++ skills.
Research-focused engineer advancing agentic model capabilities across synthetic data, task environments, evaluations, training, and usability improvements. Requires strong Python engineering, deep learning framework experience, scalable distributed training skills, and scientific experimentation ability.
Researcher focused on scaling reinforcement learning for frontier models, with ownership spanning asynchronous RL algorithms, inference and distributed training systems, and large-scale empirical studies. Requires strong Python and deep learning experience, scalable systems debugging, and rigorous research judgment.
Develop multimodal perception and authentication systems combining visual, audio, and other sensor signals for real-world AI products. The role requires machine learning expertise, practical research-to-system experience, and proficiency in Python and PyTorch with comfort in C++.