Skip to content

Staff Software Engineer - Databases SRE

Leads production reliability for Grafana Cloud’s multi-tenant database products, partnering with product engineering teams to improve SLOs, scalability, observability, automation, and incident response. Requires 8+ years of engineering experience, including substantial SRE or production engineering work, plus strong Kubernetes and cloud expertise.

About the job

Responsibilities

  • Partner closely with product engineering squads using an embedded model.
  • Own production reliability for high-SLA and complex customer environments.
  • Design and implement automation to scale reliability practices and eliminate toil.
  • Define and evolve per-tenant SLOs and reliability models.
  • Proactively reduce SLO burn and prevent repeat incidents.
  • Serve as a primary escalation point and participate in on-call rotations.
  • Lead customer-impacting incident response and post-incident reviews.
  • Contribute to design documents and code reviews.
  • Influence feature design for production scalability and operability.
  • Improve alert quality and reduce noisy escalations.
  • Improve customer observability within their environments.
  • Design and implement fault-tolerant, reliable, and scalable systems.
  • Collaborate with engineering leaders on product strategy, roadmaps, and technical designs.
  • Teach Site Reliability Engineering practices and communicate reliability best practices.
  • Investigate incidents through resolution, post-incident review, and customer communication when necessary.

Requirements

  • 8+ years of engineering experience, including 4+ years in SRE, customer reliability engineering, or production engineering.
  • Strong Kubernetes experience in AWS, Google Cloud, or Azure.
  • Familiarity with infrastructure-as-code tools such as Helm, Terraform, and Jsonnet.
  • Technical leadership experience, including leading projects and mentoring engineers.
  • Experience operating multi-tenant production systems.
  • Strong experience designing and implementing SLOs.
  • Experience with one or more programming languages, such as Go, Python, or Java.
  • Experience with Linux operating system internals.
  • Knowledge of networking, cloud storage, and scaling.
  • Excellent problem-solving and troubleshooting skills.
  • Experience participating in blame-free incident response and writing high-quality post-incident reviews.
  • Ability to reason about performance, scaling, and failure modes.
  • Ability to partner deeply with product engineering teams.
  • Comfortable working autonomously and with self-direction.
  • Intellectual curiosity, transparency, bias toward action, and a collaborative attitude.

Compensation and Benefits

  • UK base compensation range: £103,958–£124,750.
  • Actual compensation varies based on level, experience, and skillset.
  • Benefits include equity, bonus where applicable, and other benefits.
  • 30 days of annual leave per year, including three Grafana shutdown days, subject to local legislation.
  • In-person onboarding.

Skills

Kubernetes, AWS, GCP, Azure, Helm, Terraform, Jsonnet, Go, Python, Java, Linux, Networking, Cloud Storage, SLOs, Incident Response

GitLab

GitLab

Canada
Site Reliability Engineer, Intermediate to Senior Staff
$126k+/yrRemote5+ YOEDevOps / SRE

Site Reliability Engineers build and operate scalable production infrastructure, automate operational workflows, and improve observability, incident response, and service reliability. The role spans Intermediate through Senior Staff levels and requires experience with Kubernetes, infrastructure as code, cloud platforms, and software engineering.

Nango

Nango

United States
Staff Engineer, Platform & Infrastructure
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.

Nango

Nango

United States
Staff Platform Engineer
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, Observability & Profiling
£325k+/yrHybrid10+ YOEDevOps / SRE

Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.