Skip to content
AnthropicAnthropic

Staff Software Engineer, Continuous Integration

Build and operate highly reliable continuous integration infrastructure, intelligent test-selection systems, and incident-response automation at scale. The role requires 10+ years of experience with large-scale CI/CD systems, container orchestration, and developer productivity tooling.

About the job

Responsibilities

  • Design and build highly reliable, scalable CI infrastructure supporting thousands of daily builds across multiple cloud providers.
  • Develop intelligent test-selection systems that reduce CI time while maintaining code quality.
  • Build and improve incident-response automation, including cluster load shedding, automatic recovery, and observability tooling.
  • Improve test-infrastructure reliability through flake detection, quarantine systems, and test-state management.

Requirements

  • 10+ years of relevant industry experience building and operating large-scale CI/CD systems.
  • Deep experience with CI orchestration tools such as Buildkite, Jenkins, GitHub Actions, or similar.
  • Strong focus on developer productivity and reducing friction in the software development lifecycle.
  • Experience with container orchestration at scale.
  • Excellent communication skills and ability to support internal partners.
  • Strong commitment to reliability and building systems that avoid repeated failure modes.

Nice to haves

  • Experience with merge queues and branch management at scale.
  • Experience with test infrastructure, including intelligent test selection and flake management.
  • Experience building CLI tools and developer-facing services.
  • GitHub API and automation experience.

Compensation and benefits

  • Annual salary: £325,000–£390,000 GBP.
  • Bachelor's degree or equivalent combination of education, training, and experience.
  • Competitive compensation and benefits.
  • Optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours.
  • Office space for collaboration.
  • Visa sponsorship may be available depending on the role and candidate.

Skills

Continuous Integration, Continuous Delivery, Buildkite, Jenkins, GitHub Actions, Kubernetes, Github Api, Cli Tools, Test Selection, Test Infrastructure, Merge Queues, Branch Management, Observability, Cloud Infrastructure, Automation

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, Observability & Profiling
£325k+/yrHybrid10+ YOEDevOps / SRE

Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, AI Reliability Engineering
£325k+/yrHybrid7+ YOEDevOps / SRE

Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Nango

Nango

United States
Staff Engineer, Platform & Infrastructure
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.

Nango

Nango

United States
Staff Platform Engineer
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.