Skip to content
AnthropicAnthropic

Staff Software Engineer, AI Reliability Engineering

Staff software engineer responsible for improving reliability across Anthropic’s AI serving systems, from SDK and network layers through infrastructure and accelerators. The role focuses on SLOs, observability, high availability, incident response, and resilience across distributed systems.

About the job

Responsibilities

  • Develop appropriate Service Level Objectives (SLOs) for large language model serving systems, balancing availability and latency with development velocity.
  • Design and implement monitoring and observability systems across the token path.
  • Assist in designing and implementing high-availability serving infrastructure across multiple regions and cloud providers.
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
  • Support the reliability of safeguard model serving.

Requirements

  • Strong background in distributed systems, infrastructure, or reliability engineering.
  • Experience as a reliability-minded software engineer or site reliability engineer.
  • Ability to work effectively in unfamiliar systems during incidents and help drive resolution.
  • Holistic understanding of how systems compose and where system seams exist.
  • Ability to build lasting relationships and collaborate across teams.
  • Strong communication and collaboration skills.
  • Bachelor's degree or equivalent combination of education, training, and experience.

Nice-to-haves

  • Experience as an SRE, production engineer, or in a similar reliability-focused role on large-scale systems.
  • Experience operating large-scale model serving or training infrastructure involving more than 1,000 GPUs.
  • Experience with ML hardware accelerators, including GPUs, TPUs, or Trainium.
  • Understanding of ML-specific networking optimizations such as RDMA and InfiniBand.
  • Expertise in AI-specific observability tools and frameworks.
  • Experience with chaos engineering and systematic resilience testing.
  • Contributions to open-source infrastructure or ML tooling.

Compensation and Benefits

  • Annual salary: €235,000–€295,000 EUR
  • Competitive compensation and benefits.
  • Optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours.

Skills

Distributed Systems, Infrastructure, Site Reliability Engineering, Service Level Objectives, Monitoring, Observability, High Availability, Incident Response, Cloud Providers, Gpus, Tpus, Trainium, Rdma, InfiniBand, Chaos Engineering

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, Observability & Profiling
£325k+/yrHybrid10+ YOEDevOps / SRE

Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, AI Reliability Engineering
£325k+/yrHybrid7+ YOEDevOps / SRE

Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.

Nango

Nango

United States
Staff Engineer, Platform & Infrastructure
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.

Nango

Nango

United States
Staff Platform Engineer
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.