Staff Software Engineer, Observability & Profiling
Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.
About the job
Responsibilities
- Design and build scalable telemetry ingest and storage pipelines for metrics, logs, traces, and error data across multi-cluster infrastructure.
- Build low-overhead observability solutions providing deep fleet-wide visibility into system behavior.
- Own and evolve core observability platforms, including migrations and architectural improvements that improve reliability, reduce cost, and scale.
- Build instrumentation libraries, SDKs, and eBPF-based auto-instrumentation with and without code changes.
- Reduce mean time to detection and resolution through cross-signal correlation, unified query interfaces, and AI-assisted diagnostic tooling.
- Use continuous profiling and utilization telemetry to optimize CPU, memory, and accelerator fleets.
- Partner with Research, Inference, Product, and Infrastructure teams to meet their operational visibility needs.
Requirements
- Hands-on experience building and operating large-scale observability or monitoring infrastructure.
- Deep end-to-end experience with observability signals, from instrumentation through ingest, query, and analysis.
- Understanding of high-throughput telemetry pipelines and operational-data collection, storage, and querying tradeoffs at scale.
- Ability to troubleshoot below the application layer, including the kernel, network stack, or hardware.
- Excellent communication and collaboration skills.
- Bachelor’s degree or equivalent combination of education, training, and experience.
Nice-to-haves
- 10+ years of relevant industry experience, including large-scale observability or monitoring infrastructure.
- Production experience with eBPF-based tracing, profiling, or network visibility.
- Experience running continuous profiling at fleet scale, including overhead budgets and symbolization.
- Kernel- and syscall-level debugging and performance engineering experience.
- Experience profiling or instrumenting accelerator workloads.
- Experience with high-cardinality metrics systems or large-scale telemetry storage backends.
- Experience with OpenTelemetry instrumentation, collector pipelines, and tail-based sampling.
- Interest in applying AI or LLMs to root-cause analysis, anomaly detection, or intelligent alerting.
Compensation
- Annual salary: £325,000–£390,000 GBP.
- Benefits include competitive compensation, equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
Skills
Observability, Monitoring, Telemetry Pipelines, Metrics, Logging, Distributed Tracing, Continuous Profiling, Ebpf, OpenTelemetry, Kernel Debugging, Syscalls, Performance Engineering, Ai/Llms, Accelerator Workloads, Telemetry Storage
Similar jobs
DevOps / SRE jobsLeads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.
Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.
Site Reliability Engineers build and operate scalable production infrastructure, automate operational workflows, and improve observability, incident response, and service reliability. The role spans Intermediate through Senior Staff levels and requires experience with Kubernetes, infrastructure as code, cloud platforms, and software engineering.