Research, Coding Agents
Research role focused on improving agentic coding capabilities through reinforcement-learning training, synthetic data, coding environments, reward design, and evaluations. Requires strong Python engineering, scalable distributed-training experience, and a bachelor’s degree or equivalent; research experience and a PhD are preferred.
About the job
Responsibilities
- Design and run reinforcement learning training jobs targeting agentic coding capabilities, iterating on training recipes and data.
- Build and improve sandboxed coding environments and reward signals for model training and evaluation.
- Generate and curate high-quality synthetic coding data and build scalable, general-purpose data pipelines.
- Design evaluations that measure real-world coding usefulness and train models against them to improve day-to-day usability.
- Debug and analyze large-scale reinforcement learning runs to identify confounders, reward hacking, and other failure modes.
- Collaborate with infrastructure, evaluation, and post-training teams on shared data, joint training runs, and usability improvements; ship results into model releases.
Requirements
- Strong engineering skills, including the ability to contribute code and debug complex codebases.
- Proficiency in Python and familiarity with at least one deep learning framework.
- Experience debugging distributed training and writing scalable code.
- Bachelor’s degree or equivalent experience in computer science, machine learning, physics, mathematics, or a related discipline.
- Clear written communication and ability to explain complex technical concepts.
Nice-to-Haves
- Experience building synthetic data pipelines and systems adopted and maintained by a team.
- Experience identifying model-usability gaps and addressing them through custom evaluations and training data.
- Experience making large-scale agentic reinforcement learning infrastructure reliable despite failures at scale.
- Experience improving the coding capabilities of a frontier model.
- PhD in computer science, machine learning, physics, mathematics, or a related discipline, or equivalent industry research experience.
Compensation and Benefits
- Annual salary range: $350,000–$475,000 USD.
- Health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support.
- Visa sponsorship.
Skills
Python, PyTorch, TensorFlow, JAX, Reinforcement Learning, Distributed Training, Synthetic Data, Data Pipelines, Reward Modeling, Sandbox Environments, Model Evaluation, Large-Scale Training
Similar jobs
AI Research jobsResearch Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.
Conduct AI safety research across data curation, post-training, evaluations, synthetic data, and red-teaming to improve model reliability on harmful and dual-use requests. The role requires AI safety experience, Python, deep learning frameworks, and scalable technical research skills.
Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.
Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.
Conducts hands-on medicinal chemistry research to evaluate AI-generated molecules and synthetic routes, advancing small-molecule programs from design through experimental validation. Requires a chemistry PhD, sustained synthetic experience, and cross-functional collaboration skills.