Skip to content
AnthropicAnthropic

Staff Software Engineer, Environments Infrastructure

Build and own Python frameworks, APIs, and infrastructure for Anthropic's RL environments and agent runtimes. Embed with research teams to productionize their work, design for correctness in stateful distributed systems, and create self-service tooling for production debugging.

About the job

Key Responsibilities

  • Design widely used APIs, frameworks, and abstractions that other engineers and researchers build on, making correct usage the default and ruling out entire classes of errors structurally.
  • Own the platform layers that sit beneath every environment, including the agent runtime.
  • Build the tooling that lets environment owners understand, debug, and maintain their environments in production without needing an infrastructure engineer in the loop.
  • Embed with research teams on a rotational basis, work directly in their codebases without slowing down the research they support, and transfer ownership when you rotate off.
  • Anticipate silent failure modes and prevent them structurally through type safety, well-designed invariants, targeted testing, and refactors that reduce the room for correctness issues.
  • Drive adoption of new frameworks across the organization, including deprecations and cutovers.
  • Help define the engineering standards, review practices, and design patterns for a new team, and mentor researchers and engineers in adopting them.

Minimum Qualifications

  • Deep expertise in Python, including static typing, safe async and concurrency patterns, and writing performant code.
  • Strong taste in API and framework design, the ability to explain why an interface is right or wrong rather than just recognizing it, and a track record of other engineers or teams adopting and building on frameworks you have built.
  • Experience designing or operating stateful concurrent or distributed systems, and reasoning carefully about failure, retries, idempotency, and consistency.
  • A habit of verification: you measure before you conclude, and you build the checks that let a system show it's correct.
  • Experience working productively in large, evolving, or research-style codebases that you didn't originally write.
  • Strong written and verbal communication with collaborators of varied engineering backgrounds, and comfort with ambiguity: able to scope your own work from a loosely defined problem and drive it to a maintainable outcome.

Preferred Qualifications

  • Experience building infrastructure, tooling, or frameworks for machine learning research or RL workflows, and familiarity with agentic systems or LLM training pipelines.
  • Experience building agent frameworks, orchestration engines, or multi-agent systems, including checkpoint and restore, replay, and coordination of long-running stateful processes.
  • Experience using AI coding tools on code where correctness matters, with good judgment about what to delegate and how to make the results verifiable.
  • Experience building client libraries or SDKs on top of sandboxed, containerized, or remote execution platforms.
  • Experience with large-scale data processing, dataset lifecycle management, or data lineage systems.
  • Experience designing serialization schemes, plugin systems, or extensible class hierarchies used across an organization.
  • Experience embedding with or consulting for other teams and handing off systems for others to own, or defining code standards adopted across teams, or prior experience as a technical lead.

Representative Projects

  • Design a base RL environment abstraction that can be subclassed to support the large majority of environments built across RL.
  • Redesign the model-tool interface for sandboxed agentic environments so that state is guaranteed to survive serialization, making it structurally impossible to write a tool that silently loses state.
  • Design the state-sharing and recovery model for multi-agent workloads, so that losing a sandbox partway through a task becomes a transparent resume rather than lost work.
  • Define the failure and retry model for a sandboxed execution platform, distinguishing infrastructure faults from genuine task outcomes so that each is handled correctly.
  • Build the tooling that lets an environment owner diagnose why their environment is unhealthy in a production run, and fix it themselves.

Annual Salary: $405000—$605000 USD

Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience.

Skills

Python, Static Typing, Async, Concurrency, API Design, Framework Design, Distributed Systems, Stateful Systems, Machine Learning, Reinforcement Learning, Agent Frameworks, Orchestration Engines, Multi-Agent Systems, Serialization, Data Processing

Anthropic

Anthropic

San Francisco, CA

Staff+ Software Engineer, ML Inference Path
$320k+/yrHybrid7+ YOEML Engineering

Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.

Garner Health

Garner Health

New York, NY

Staff Applied Scientist
$300k+/yrHybrid7+ YOEML Engineering

Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.

Garner Health

Garner Health

New York, NY

Staff Machine Learning Operations Engineer
$298k+/yrHybrid7+ YOEML Engineering

Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.

Reddit

Reddit

United States

Senior Staff Machine Learning Systems Engineer, Ads ML Platform
$293k+/yrRemote8+ YOEML Engineering

Leads technical strategy for Reddit’s Ads ML Platform, improving feature development, training-data generation, experimentation, and the path to production ML serving. The role requires 8+ years in infrastructure or distributed systems, production ML platform experience, and strong cross-team technical leadership.

Square

Square

San Francisco, CA

Staff Machine Learning Engineer, Fraud & Abuse
$277k+/yrRemote12+ YOEML Engineering

Build and operate production machine learning systems for ranking, retrieval, recommendations, personalization, and customer intelligence. The role requires 12+ years of production software and ML experience, strong expertise in intelligent systems, and sound judgment around trustworthy customer-impacting signals.