# Staff Software Engineer, Code RL

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** San Francisco, CA, New York, NY, Seattle, WA
**Role:** ML Engineering
**Salary:** $405k – $625k/yr
**Experience:** 7+ years
**Skills:** Python, Static Typing, Async Patterns, Concurrency, API Design, Framework Design, Large Codebases, System Design, Testing, Reinforcement Learning, Distributed Systems, Machine Learning Infrastructure, Llm Training, Data Processing, Lint Rules
**Posted:** 2026-07-25

> Staff Software Engineer driving reinforcement learning infrastructure for Claude's coding capabilities at Anthropic. Design APIs/frameworks, embed with research teams to build and hand off maintainable systems, improve research code reliability, and ensure production RL run health. Requires deep Python expertise, API design track record, and failure-mode intuition.

## Job Description

## Key Responsibilities
- Design widely-used APIs, frameworks, and abstractions that other engineers and researchers build on, with careful attention to interface legibility and principled defaults.
- Embed with research teams on a rotational basis: understand their engineering needs, build systems and APIs that support their work, and transfer ownership so teams can maintain those systems after you rotate off.
- Work directly in research codebases, improving reliability and structure without slowing down the research they support.
- Anticipate silent failure modes and prevent them structurally through type safety, well-designed invariants, targeted testing, and refactors that shrink the surface area for bugs.
- Contribute to the reliability of production RL systems, including monitoring, regression detection, and triage tooling.
- Help define engineering standards, review practices, and design patterns for a new team, and mentor researchers and engineers in adopting them.

## Minimum Qualifications
- Deep expertise in Python, including static typing, safe async and concurrency patterns, and writing performant Python code.
- A track record of designing intuitive, safe APIs or frameworks that other engineers or teams adopted and built on.
- Experience working productively in large, evolving, or research-style codebases that you didn't originally write.
- Demonstrated ability to anticipate failure modes — especially silent ones — and prevent them structurally through system design, type safety, and testing.
- Strong written and verbal communication skills, including the ability to explain system designs to collaborators with varied engineering backgrounds.
- Comfort with ambiguity: able to scope your own work from a loosely defined problem and drive it to a maintainable outcome.

## Preferred Qualifications
- Experience building infrastructure, tooling, or frameworks for machine learning research or RL workflows.
- Familiarity with reinforcement learning concepts, agentic systems, or LLM training pipelines.
- Experience building or operating large-scale distributed systems.
- Experience building client libraries or SDKs on top of sandboxed, containerized, or remote execution platforms.
- Experience with large-scale data processing or dataset lifecycle management.
- Experience designing plugin systems or extensible class hierarchies used across an organization.
- Experience embedding with or consulting for other teams, including successfully handing off systems for others to own.
- Experience defining code standards, lint rules, or static verification approaches adopted across multiple teams.
- Prior experience as a technical lead, or setting engineering standards for a team.
- Prior experience maintaining an open source project.

## Representative Projects
- Design a base RL environment abstraction general enough to be subclassed across a wide range of environments.
- Design a model-tool interface for sandboxed agentic environments that has explicit serialization semantics.
- Partner with the platform teams that own the sandbox runtime to specify low-level features that improve the integrity of agentic coding tasks.
- Design probes that catch sandbox regressions early.
- Design the lifecycle and maintenance scheme for a production dataset.
- Lead a research code refactor replacing loosely structured data containers with equivalents that carry stronger correctness guarantees, without breaking the experiments that depend on them.
- Design lint rules and code-style requirements that favor statically verifiable patterns — including patterns less likely to be overlooked by an LLM reviewing or writing the code — to shrink the surface area for silent bugs.
- Build the access layer that lets researchers discover and reuse data artifacts across teams.

**Annual Salary:** $405000—$625000 USD

**Minimum education:** Bachelor’s degree or an equivalent combination of education, training, and/or experience.

## Similar jobs

- [Staff+ Software Engineer, ML Inference Path](https://hotfix.jobs/jobs/90cf7427-87de-41e1-af88-0d4e1187b94b) - Anthropic - San Francisco, CA - $320k – $485k/yr
- [Staff Applied Scientist](https://hotfix.jobs/jobs/d518c7a7-07d4-4f06-af79-9c478df769ca) - Garner Health - New York, NY - $300k – $390k/yr
- [Staff Machine Learning Operations Engineer](https://hotfix.jobs/jobs/b5d46280-8221-4b6e-b0b2-033cc6f2c8d7) - Garner Health - New York, NY - $298k – $351k/yr
- [Senior Staff Machine Learning Systems Engineer, Ads ML Platform](https://hotfix.jobs/jobs/efc443bf-db64-4fa3-b123-eb399c454a93) - Reddit - Remote - $293k – $410k/yr
- [Staff Machine Learning Engineer, Fraud & Abuse](https://hotfix.jobs/jobs/2068a66a-3fd2-42f6-ab53-d199be305035) - Square - Remote - $277k – $415k/yr

**Apply:** https://hotfix.jobs/jobs/35ad5b7b-cb43-4bb9-9a06-9e5091c38bd3
**Canonical:** https://hotfix.jobs/jobs/35ad5b7b-cb43-4bb9-9a06-9e5091c38bd3