Staff+ Site Reliability Engineer, Safeguards ML Infra
Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
About the job
Responsibilities
- Configure, deploy, and verify safeguards for every new model release.
- Serve as the safeguards point of contact during release windows and launch operations.
- Deploy new safety classifiers from research, including canary rollouts, post-deployment validation, and discrepancy investigation.
- Verify safeguards across first-party infrastructure, AWS Bedrock, and Google Cloud Vertex AI, and eliminate configuration drift.
- Automate launch runbooks, manual checks, and one-off deployments into continuous validation and repeatable pipelines.
- Build and maintain a safeguards registry with provenance for production systems, models, platforms, deployment times, and deployers.
- Participate in on-call and operational rotations covering incidents, model provisioning, and time-sensitive launches.
- Use postmortems to drive process and tooling improvements.
Requirements
- Experience owning production change management at scale, including deployment pipelines, configuration management, or canary analysis.
- Experience with high-stakes releases as a launch captain, incident commander, or release owner.
- Meaningful production on-call and incident-response experience.
- Experience reducing operational toil through automation and transitioning manual deployment processes to self-service pipelines.
- Experience operating cloud platforms at scale, particularly AWS and Google Cloud.
- Proficiency in Python.
- Bachelor's degree or equivalent combination of education, training, and experience.
- 8+ years of industry software engineering or site reliability engineering experience.
Nice to Have
- Rust experience.
- Experience running launch or production-readiness reviews across multiple teams.
- Familiarity with LLM inference systems and transformer-based model operations.
Compensation
- Annual salary: $320,000–$485,000 USD.
Skills
Python, Rust, AWS, GCP, Aws Bedrock, Google Cloud Vertex Ai, Deployment Pipelines, Configuration Management, Canary Analysis, Incident Response, On-Call Operations, Llm Inference
Similar jobs
DevOps / SRE jobsOwn the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.
Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.