Skip to content
AnthropicAnthropic

Staff Software Engineer, AI Reliability Engineering

Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.

About the job

Responsibilities

  • Develop Service Level Objectives for large language model serving systems, balancing availability, latency, and development velocity.
  • Design and implement monitoring and observability systems across the token path.
  • Help design and implement highly available serving infrastructure across multiple regions and cloud providers.
  • Lead incident response for critical AI services, driving rapid recovery, thorough incident reviews, and systematic improvements.
  • Support the reliability of safeguard model serving.

Requirements

  • Strong background in distributed systems, infrastructure, or reliability engineering.
  • Reliability-minded software engineering or Site Reliability Engineering experience.
  • Ability to troubleshoot unfamiliar systems during incidents and drive resolution.
  • Holistic understanding of system composition and integration points.
  • Strong communication, collaboration, relationship-building, and ownership skills.

Nice-to-haves

  • Experience as an SRE, Production Engineer, or in a similar reliability-focused role on large-scale systems.
  • Experience operating large-scale model serving or training infrastructure with more than 1,000 GPUs.
  • Experience with ML hardware accelerators, including GPUs, TPUs, or Trainium.
  • Knowledge of RDMA and InfiniBand.
  • Expertise with AI-specific observability tools and frameworks.
  • Experience with chaos engineering and systematic resilience testing.
  • Contributions to open-source infrastructure or ML tooling.

Compensation and Benefits

  • Annual salary: £325,000–£390,000 GBP.
  • Benefits include competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration support.
  • Bachelor’s degree or an equivalent combination of education, training, and experience is required.
  • Field of study should be relevant to the role through coursework, training, or professional experience.

Skills

Distributed Systems, Infrastructure, Site Reliability Engineering, Service Level Objectives, Observability, Monitoring, High Availability, Cloud Computing, Incident Response, GPU, Tpu, Trainium, Rdma, InfiniBand, Chaos Engineering

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, Observability & Profiling
£325k+/yrHybrid10+ YOEDevOps / SRE

Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Nango

Nango

United States
Staff Engineer, Platform & Infrastructure
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.

Nango

Nango

United States
Staff Platform Engineer
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.

GitLab

GitLab

Canada
Site Reliability Engineer, Intermediate to Senior Staff
$126k+/yrRemote5+ YOEDevOps / SRE

Site Reliability Engineers build and operate scalable production infrastructure, automate operational workflows, and improve observability, incident response, and service reliability. The role spans Intermediate through Senior Staff levels and requires experience with Kubernetes, infrastructure as code, cloud platforms, and software engineering.