# Staff+ Site Reliability Engineer, Safeguards ML Infra

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** San Francisco, CA, Seattle, WA, New York, NY
**Role:** DevOps / SRE
**Salary:** $320k – $485k/yr
**Experience:** 8+ years
**Skills:** Python, Rust, AWS, GCP, Aws Bedrock, Google Cloud Vertex Ai, Deployment Pipelines, Configuration Management, Canary Analysis, Incident Response, On-Call Operations, Llm Inference
**Posted:** 2026-09-04

> Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.

## Job Description

## Responsibilities
- Configure, deploy, and verify safeguards for every new model release.
- Serve as the safeguards point of contact during release windows and launch operations.
- Deploy new safety classifiers from research, including canary rollouts, post-deployment validation, and discrepancy investigation.
- Verify safeguards across first-party infrastructure, AWS Bedrock, and Google Cloud Vertex AI, and eliminate configuration drift.
- Automate launch runbooks, manual checks, and one-off deployments into continuous validation and repeatable pipelines.
- Build and maintain a safeguards registry with provenance for production systems, models, platforms, deployment times, and deployers.
- Participate in on-call and operational rotations covering incidents, model provisioning, and time-sensitive launches.
- Use postmortems to drive process and tooling improvements.

## Requirements
- Experience owning production change management at scale, including deployment pipelines, configuration management, or canary analysis.
- Experience with high-stakes releases as a launch captain, incident commander, or release owner.
- Meaningful production on-call and incident-response experience.
- Experience reducing operational toil through automation and transitioning manual deployment processes to self-service pipelines.
- Experience operating cloud platforms at scale, particularly AWS and Google Cloud.
- Proficiency in Python.
- Bachelor's degree or equivalent combination of education, training, and experience.
- 8+ years of industry software engineering or site reliability engineering experience.

## Nice to Have
- Rust experience.
- Experience running launch or production-readiness reviews across multiple teams.
- Familiarity with LLM inference systems and transformer-based model operations.

## Compensation
- Annual salary: $320,000–$485,000 USD.

## Similar jobs

- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/b1ad3336-2ce1-4ede-a90f-901a178fe108) - Headway - Remote - $265k – $331k/yr
- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/6e086905-4170-407a-a7a4-2e05df0701c5) - Polymarket - New York, NY - $250k – $500k/yr
- [Senior Staff Deployment Automation Engineer](https://hotfix.jobs/jobs/e7c5a05a-7667-45ba-8f89-00349f1e9aaa) - Crusoe - San Francisco, CA - $250k – $300k/yr
- [Senior Staff Software Engineer, DC Infrastructure](https://hotfix.jobs/jobs/5e51b5cc-872b-4556-a065-f5f95a363fad) - Crusoe - San Francisco, CA - $250k – $300k/yr
- [Staff Engineer - Cloud Networks](https://hotfix.jobs/jobs/66aff9c2-57a3-4f6b-bc39-51a61a03c8ae) - Datadog - Boston, MA - $244k – $305k/yr

**Apply:** https://hotfix.jobs/jobs/6492550a-2ff4-498b-8247-470adae7d0c3
**Canonical:** https://hotfix.jobs/jobs/6492550a-2ff4-498b-8247-470adae7d0c3