Skip to content
AnthropicAnthropic

Staff Software Engineer, Observability & Profiling

Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.

About the job

Responsibilities

  • Design and build scalable telemetry ingest and storage pipelines for metrics, logs, traces, and error data across multi-cluster infrastructure.
  • Build low-overhead observability solutions providing deep fleet-wide visibility into system behavior.
  • Own and evolve core observability platforms, including migrations and architectural improvements that improve reliability, reduce cost, and scale.
  • Build instrumentation libraries, SDKs, and eBPF-based auto-instrumentation with and without code changes.
  • Reduce mean time to detection and resolution through cross-signal correlation, unified query interfaces, and AI-assisted diagnostic tooling.
  • Use continuous profiling and utilization telemetry to optimize CPU, memory, and accelerator fleets.
  • Partner with Research, Inference, Product, and Infrastructure teams to meet their operational visibility needs.

Requirements

  • Hands-on experience building and operating large-scale observability or monitoring infrastructure.
  • Deep end-to-end experience with observability signals, from instrumentation through ingest, query, and analysis.
  • Understanding of high-throughput telemetry pipelines and operational-data collection, storage, and querying tradeoffs at scale.
  • Ability to troubleshoot below the application layer, including the kernel, network stack, or hardware.
  • Excellent communication and collaboration skills.
  • Bachelor’s degree or equivalent combination of education, training, and experience.

Nice-to-haves

  • 10+ years of relevant industry experience, including large-scale observability or monitoring infrastructure.
  • Production experience with eBPF-based tracing, profiling, or network visibility.
  • Experience running continuous profiling at fleet scale, including overhead budgets and symbolization.
  • Kernel- and syscall-level debugging and performance engineering experience.
  • Experience profiling or instrumenting accelerator workloads.
  • Experience with high-cardinality metrics systems or large-scale telemetry storage backends.
  • Experience with OpenTelemetry instrumentation, collector pipelines, and tail-based sampling.
  • Interest in applying AI or LLMs to root-cause analysis, anomaly detection, or intelligent alerting.

Compensation

  • Annual salary: £325,000–£390,000 GBP.
  • Benefits include competitive compensation, equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.

Skills

Observability, Monitoring, Telemetry Pipelines, Metrics, Logging, Distributed Tracing, Continuous Profiling, Ebpf, OpenTelemetry, Kernel Debugging, Syscalls, Performance Engineering, Ai/Llms, Accelerator Workloads, Telemetry Storage

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, AI Reliability Engineering
£325k+/yrHybrid7+ YOEDevOps / SRE

Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Nango

Nango

United States
Staff Engineer, Platform & Infrastructure
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.

Nango

Nango

United States
Staff Platform Engineer
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.

GitLab

GitLab

Canada
Site Reliability Engineer, Intermediate to Senior Staff
$126k+/yrRemote5+ YOEDevOps / SRE

Site Reliability Engineers build and operate scalable production infrastructure, automate operational workflows, and improve observability, incident response, and service reliability. The role spans Intermediate through Senior Staff levels and requires experience with Kubernetes, infrastructure as code, cloud platforms, and software engineering.