Skip to content
AnthropicAnthropic

Staff+ Site Reliability Engineer, Safeguards ML Infra

Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.

About the job

Responsibilities

  • Configure, deploy, and verify safeguards for every new model release.
  • Serve as the safeguards point of contact during release windows and launch operations.
  • Deploy new safety classifiers from research, including canary rollouts, post-deployment validation, and discrepancy investigation.
  • Verify safeguards across first-party infrastructure, AWS Bedrock, and Google Cloud Vertex AI, and eliminate configuration drift.
  • Automate launch runbooks, manual checks, and one-off deployments into continuous validation and repeatable pipelines.
  • Build and maintain a safeguards registry with provenance for production systems, models, platforms, deployment times, and deployers.
  • Participate in on-call and operational rotations covering incidents, model provisioning, and time-sensitive launches.
  • Use postmortems to drive process and tooling improvements.

Requirements

  • Experience owning production change management at scale, including deployment pipelines, configuration management, or canary analysis.
  • Experience with high-stakes releases as a launch captain, incident commander, or release owner.
  • Meaningful production on-call and incident-response experience.
  • Experience reducing operational toil through automation and transitioning manual deployment processes to self-service pipelines.
  • Experience operating cloud platforms at scale, particularly AWS and Google Cloud.
  • Proficiency in Python.
  • Bachelor's degree or equivalent combination of education, training, and experience.
  • 8+ years of industry software engineering or site reliability engineering experience.

Nice to Have

  • Rust experience.
  • Experience running launch or production-readiness reviews across multiple teams.
  • Familiarity with LLM inference systems and transformer-based model operations.

Compensation

  • Annual salary: $320,000–$485,000 USD.

Skills

Python, Rust, AWS, GCP, Aws Bedrock, Google Cloud Vertex Ai, Deployment Pipelines, Configuration Management, Canary Analysis, Incident Response, On-Call Operations, Llm Inference

Headway

Headway

San Francisco, CA
Staff Infrastructure Engineer
$265k+/yrRemote8+ YOEDevOps / SRE

Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Deployment Automation Engineer
$250k+/yrOn-site12+ YOEDevOps / SRE

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Software Engineer, DC Infrastructure
$250k+/yrOn-site7+ YOEDevOps / SRE

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.

Datadog

Datadog

Boston, MA
Staff Engineer - Cloud Networks
$244k+/yrHybrid7+ YOEDevOps / SRE

Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.