Skip to content
AnthropicAnthropic

Research Engineer, Code RL

Research Engineer advancing Claude's code generation capabilities through reinforcement learning. Design RL environments, build verifiers, run training experiments on frontier models, and improve training pipelines for real software engineering tasks.

About the job

Responsibilities

  • Design RL environments and coding tasks for training models on real software engineering work
  • Build reward signals and verifiers that capture what "good code" means
  • Run training experiments on frontier models
  • Diagnose why models do or don't improve at classes of software-engineering work
  • Improve speed and reliability of training pipelines
  • Advance models' ability to write, edit, test, debug, and ship real software end-to-end

Requirements

  • Strong software-engineering skills and deep Python expertise, including async/concurrent programming
  • Comfortable owning systems end to end and debugging across the stack
  • Ability to balance research exploration with engineering implementation
  • Rigorous approach to experimental design and interpreting results
  • Care about code quality, testing, and performance
  • Commitment to developing safe and beneficial AI systems
  • Bachelor's degree or equivalent combination of education, training, and/or experience in a relevant field

Nice-to-Haves

  • Experience with reinforcement learning, RLHF, post-training, or LLM finetuning
  • Experience building coding agents, code-execution sandboxes, eval harnesses, verifiers, or developer tooling
  • Background in program analysis, testing, verification, compilers, or formal methods
  • Experience with PyTorch and large-scale distributed training; performance profiling and optimization of ML systems
  • CUDA / GPU or TPU kernel experience and accelerator-performance intuition
  • Experience with virtualization and sandboxed code execution environments

Compensation & Benefits

  • Annual Salary: $500,000—$850,000 USD
  • Competitive compensation and benefits
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours
  • Visa sponsorship available

Skills

Python, Reinforcement Learning, PyTorch, CUDA, Gpu Programming, Distributed Training, Async Programming, Code Execution Sandboxes, Program Analysis, Formal Methods

Garner Health

Garner Health

New York, NY

Manager, Applied Science
$300k+/yrHybrid8+ YOEML Engineering

Leads and builds a team of Applied Scientists developing production algorithmic systems for healthcare optimization, LLM applications, and member engagement. Requires 6+ years of relevant industry experience, strong technical judgment, and hands-on expertise across machine learning and optimization.

OpenAI

OpenAI

San Francisco, CA

Machine Learning Engineer, API Multicloud
$295k+/yrOn-site7+ YOEML Engineering

Build production machine learning systems for model customization, post-training, evaluation, and AWS-native API integration. The role requires 7+ years of relevant engineering experience and expertise in deep learning, transformers, LLM fine-tuning, and production ML infrastructure.

OpusClip

OpusClip

Mountain View, CA

AI Engineering Lead
$280k+/yrOn-siteML Engineering

Leads a hands-on AI engineering team developing, evaluating, and deploying large-scale multimodal and video models. The role combines post-training, inference optimization, product experimentation, technical roadmap ownership, and people management.

Discord

Discord

United States

Engineering Manager, Machine Learning
$272k+/yrOn-site8+ YOEML Engineering

Leads Discord’s Safety ML team, setting technical direction and overseeing production machine learning systems for content understanding, account integrity, and platform abuse. Requires substantial machine learning and engineering management experience, hands-on technical depth, and experience delivering ML systems at scale.

Baselayer

Baselayer

San Francisco, CA

Senior AI Engineer, Agentic Data Enrichment
$230k+/yrHybrid5+ YOEML Engineering

Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.