# Staff Software Engineer, AI Reliability Engineering

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** London, United Kingdom
**Role:** DevOps / SRE
**Salary:** £325k – £390k/yr
**Experience:** 7+ years
**Skills:** Distributed Systems, Infrastructure, Site Reliability Engineering, Service Level Objectives, Observability, Monitoring, High Availability, Cloud Computing, Incident Response, GPU, Tpu, Trainium, Rdma, InfiniBand, Chaos Engineering
**Posted:** 2026-08-28

> Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.

## Job Description

## Responsibilities
- Develop Service Level Objectives for large language model serving systems, balancing availability, latency, and development velocity.
- Design and implement monitoring and observability systems across the token path.
- Help design and implement highly available serving infrastructure across multiple regions and cloud providers.
- Lead incident response for critical AI services, driving rapid recovery, thorough incident reviews, and systematic improvements.
- Support the reliability of safeguard model serving.

## Requirements
- Strong background in distributed systems, infrastructure, or reliability engineering.
- Reliability-minded software engineering or Site Reliability Engineering experience.
- Ability to troubleshoot unfamiliar systems during incidents and drive resolution.
- Holistic understanding of system composition and integration points.
- Strong communication, collaboration, relationship-building, and ownership skills.

## Nice-to-haves
- Experience as an SRE, Production Engineer, or in a similar reliability-focused role on large-scale systems.
- Experience operating large-scale model serving or training infrastructure with more than 1,000 GPUs.
- Experience with ML hardware accelerators, including GPUs, TPUs, or Trainium.
- Knowledge of RDMA and InfiniBand.
- Expertise with AI-specific observability tools and frameworks.
- Experience with chaos engineering and systematic resilience testing.
- Contributions to open-source infrastructure or ML tooling.

## Compensation and Benefits
- Annual salary: £325,000–£390,000 GBP.
- Benefits include competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration support.
- Bachelor’s degree or an equivalent combination of education, training, and experience is required.
- Field of study should be relevant to the role through coursework, training, or professional experience.

## Similar jobs

- [Staff Software Engineer, Observability & Profiling](https://hotfix.jobs/jobs/be415182-687b-4ae9-a1c2-4b70b877514b) - Anthropic - London, United Kingdom - £325k – £390k/yr
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr
- [Staff Engineer, Platform & Infrastructure](https://hotfix.jobs/jobs/48b953ba-a6fb-4833-9ad0-140dfe4943dc) - Nango - Remote - $140k – $220k/yr
- [Staff Platform Engineer](https://hotfix.jobs/jobs/37cd4cd2-d007-4013-8153-5443ae70f1dd) - Nango - Remote - $140k – $220k/yr
- [Site Reliability Engineer, Intermediate to Senior Staff](https://hotfix.jobs/jobs/cb4f85a5-5a0b-4be9-a1ab-cd4b0c01d37b) - GitLab - Remote - $126k – $314k/yr

**Apply:** https://hotfix.jobs/jobs/66b246fd-6b0e-4283-b1ff-c460780b61a2
**Canonical:** https://hotfix.jobs/jobs/66b246fd-6b0e-4283-b1ff-c460780b61a2