Skip to content
AnthropicAnthropic

[Pipeline] Staff+ Software Engineer, Developer Acceleration

Staff+ Software Engineer building Anthropic's Agent Runtime Platform and knowledge infrastructure to enable thousands of employees to be highly productive with AI agents. Requires 10+ years large-scale distributed systems experience and agent expertise to define agentic productivity, build runtimes, write evals, and drive operational excellence.

About the job

Responsibilities

  • Define the future of agentic productivity at scale and at the frontier
  • Own the technical strategy and roadmap for your area, translating goals into concrete execution
  • Define what quality means for agent-driven engineering work, and hold that line across the company
  • Stay hands-on: build and ship the runtime platform and tooling
  • Build agent harnesses and experiment with various context management strategies to improve agent performance and correctness
  • Write evals to benchmark agent behaviors
  • Deliver impact by collaborating cross-team
  • Own infrastructure scalability and reliability, and establish operational excellence practices

Requirements

  • 10+ years building and operating large-scale distributed systems
  • 3+ years of experience leading large scale, complex projects or teams as an engineer or tech lead
  • Experience with agents, e.g. agents doing technical/knowledge work
  • Obsessed with productivity and transforming how we work
  • Experience building scalable platforms
  • Excellent communication skills and enjoy supporting internal partners
  • Bachelor’s degree or an equivalent combination of education, training, and/or experience in a relevant field

Nice-to-Haves

  • Deep understanding of, and care for, how human work will evolve; and a drive to help shepherd that transition
  • Experience with container or VM orchestration at scale
  • Experience working with researchers and engineers
  • Developer productivity or infrastructure experience, such as CI/CD, builds, etc.
  • Experience building widely adopted CLI tools and services

Skills

Distributed Systems, Agentic Systems, Agent Runtime Platforms, Context Management, Evals, Scalable Platforms, Container Orchestration, Vm Orchestration, CI/CD, Cli Tools

Anthropic

Anthropic

San Francisco, CA
Staff+ Software Engineer, Platform Portability
$405k+/yrHybrid8+ YOEDevOps / SRE

Build and operate portable infrastructure that enables Claude to run reliably across multiple cloud providers and accelerator platforms. The role requires 8+ years of distributed-systems experience, multi-cloud architecture expertise, production programming, Kubernetes, and Infrastructure as Code proficiency.

Anthropic

Anthropic

San Francisco, CA
Staff+ Site Reliability Engineer, Safeguards ML Infra
$320k+/yrHybrid8+ YOEDevOps / SRE

Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.

Headway

Headway

San Francisco, CA
Staff Infrastructure Engineer
$265k+/yrRemote8+ YOEDevOps / SRE

Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Deployment Automation Engineer
$250k+/yrOn-site12+ YOEDevOps / SRE

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.