Research Engineer, RL Scaling Science
Conduct large-scale reinforcement learning experiments, develop long-horizon benchmarks, investigate scaling behavior, and translate validated research into production training recipes. The role requires strong empirical research skills, Python, distributed ML experience, and a bachelor's degree or equivalent experience.
About the job
Responsibilities
- Design, run, and interpret large-scale reinforcement learning experiments, reasoning rigorously about what the data does and does not show.
- Investigate how reinforcement learning improves as horizon, compute, and model size grow.
- Build and maintain benchmarks for long-horizon reinforcement learning so progress is measurable and reproducible.
- Translate validated findings into production training recipes, exercising judgment about when a result is robust enough to ship.
- Debug complex issues at the seam between research and infrastructure, including failures that appear only at scale.
- Partner closely with adjacent reinforcement learning teams across research and engineering and advance the overall reinforcement learning stack.
Requirements
- Strong empirical research skills in reinforcement learning, large-scale machine learning training, or a closely adjacent area.
- Demonstrated ability to own large experiments end to end, from design through interpretation.
- Proficiency in Python and experience working with large-scale or distributed machine learning systems.
- Comfort operating at the research and systems boundary, including debugging where the two meet.
- Care about the societal impacts of artificial intelligence and responsible scaling.
- Bachelor's degree or an equivalent combination of education, training, and experience in a relevant field, as demonstrated through coursework, training, or professional experience.
Nice-to-haves
- Published or shipped work in long-horizon reinforcement learning or reinforcement learning fundamentals.
- Experience translating research findings into production training recipes.
- Demonstrated large-scale industry impact through reinforcement learning interventions.
- Experience working on frontier-scale training runs with long trajectories.
Representative Projects
- Design a benchmark suite for long-horizon reinforcement learning that distinguishes genuine capability gains from artifacts of evaluation setup.
- Take a promising experimental finding, stress-test it across model scales, and work with training teams to land it in a production recipe.
- Investigate an unexpected scaling trend in a reinforcement learning run and trace it to a root cause spanning algorithm, data, and infrastructure.
Compensation and Benefits
- Annual salary: £375,000–£640,000 GBP.
- Hybrid policy: staff are expected to work from an office at least 25% of the time; some roles may require more office time.
- Visa sponsorship is available for eligible roles and candidates.
- Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
Skills
Reinforcement Learning, Python, Machine Learning, Distributed Systems, Large-Scale Training, Long-Horizon Rl, Benchmarking, Experiment Design, Production Training, Debugging, Scaling Laws
Similar jobs
ML Engineering jobsResearch Engineer responsible for operating and improving large-scale production pretraining systems, from performance optimization and hardware debugging to experiments, observability, and launch incident response. Requires deep ML systems expertise and experience with LLM training, JAX, TPU, PyTorch, or distributed systems.
Build and ship production agentic AI workflows for complex real estate and built-world processes. The role combines product engineering, applied AI, customer collaboration, workflow orchestration, evaluation, and reliable user-facing experiences.
Build and operate Dougie, an agentic AI system that executes workflows, evaluates its own performance, retains institutional context, and improves in production. The role requires experience deploying unattended agentic systems and engineering reliable memory, retrieval, orchestration, and feedback loops.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and deploy AI-powered products for digital-native customers, taking systems from experimentation through production and scale. The role requires strong Python skills, hands-on production engineering, systematic AI evaluation, and the ability to navigate reliability, security, governance, and customer impact.