Research Engineer, Code RL
Research Engineer advancing Claude's code generation capabilities through reinforcement learning. Design RL environments, build verifiers, run training experiments on frontier models, and improve training pipelines for real software engineering tasks.
About the job
Responsibilities
- Design RL environments and coding tasks for training models on real software engineering work
- Build reward signals and verifiers that capture what "good code" means
- Run training experiments on frontier models
- Diagnose why models do or don't improve at classes of software-engineering work
- Improve speed and reliability of training pipelines
- Advance models' ability to write, edit, test, debug, and ship real software end-to-end
Requirements
- Strong software-engineering skills and deep Python expertise, including async/concurrent programming
- Comfortable owning systems end to end and debugging across the stack
- Ability to balance research exploration with engineering implementation
- Rigorous approach to experimental design and interpreting results
- Care about code quality, testing, and performance
- Commitment to developing safe and beneficial AI systems
- Bachelor's degree or equivalent combination of education, training, and/or experience in a relevant field
Nice-to-Haves
- Experience with reinforcement learning, RLHF, post-training, or LLM finetuning
- Experience building coding agents, code-execution sandboxes, eval harnesses, verifiers, or developer tooling
- Background in program analysis, testing, verification, compilers, or formal methods
- Experience with PyTorch and large-scale distributed training; performance profiling and optimization of ML systems
- CUDA / GPU or TPU kernel experience and accelerator-performance intuition
- Experience with virtualization and sandboxed code execution environments
Compensation & Benefits
- Annual Salary: $500,000—$850,000 USD
- Competitive compensation and benefits
- Optional equity donation matching
- Generous vacation and parental leave
- Flexible working hours
- Visa sponsorship available
Skills
Python, Reinforcement Learning, PyTorch, CUDA, Gpu Programming, Distributed Training, Async Programming, Code Execution Sandboxes, Program Analysis, Formal Methods
Similar jobs
ML Engineering jobsLeads and builds a team of Applied Scientists developing production algorithmic systems for healthcare optimization, LLM applications, and member engagement. Requires 6+ years of relevant industry experience, strong technical judgment, and hands-on expertise across machine learning and optimization.
Build production machine learning systems for model customization, post-training, evaluation, and AWS-native API integration. The role requires 7+ years of relevant engineering experience and expertise in deep learning, transformers, LLM fine-tuning, and production ML infrastructure.
Leads a hands-on AI engineering team developing, evaluating, and deploying large-scale multimodal and video models. The role combines post-training, inference optimization, product experimentation, technical roadmap ownership, and people management.
Leads Discord’s Safety ML team, setting technical direction and overseeing production machine learning systems for content understanding, account integrity, and platform abuse. Requires substantial machine learning and engineering management experience, hands-on technical depth, and experience delivering ML systems at scale.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.