Skip to content
GitLabGitLab

Site Reliability Engineer, Intermediate to Senior Staff

Site Reliability Engineers build and operate scalable production infrastructure, automate operational workflows, and improve observability, incident response, and service reliability. The role spans Intermediate through Senior Staff levels and requires experience with Kubernetes, infrastructure as code, cloud platforms, and software engineering.

About the job

What You'll Do

  • Keep user-facing services and production systems reliable, scalable, and efficient.
  • Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows.
  • Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling.
  • Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps.
  • Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately.
  • Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early.
  • Take part in incident response and post-incident reviews, turning learnings into changes in automation and process.
  • Document runbooks, architecture decisions, and reviews so findings become repeatable practices.

What You'll Bring

  • Experience keeping production systems reliable, combining an operations mindset with software engineering practice.
  • Experience building net-new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators or controllers, and production automation or services written from scratch.
  • Ability to read, debug, and reason about code, including behavior, performance, and failure modes. Most teams work in Go; some work in Ruby.
  • Experience with infrastructure as code, Kubernetes, and its ecosystem.
  • Hands-on experience with at least one major cloud provider, such as Google Cloud or AWS.
  • Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs.
  • Comfort participating in on-call and incident response, with a structured approach to troubleshooting under pressure.
  • Strong written communication and ability to operate independently in an async, distributed environment.
  • Track record of using automation and AI to reduce toil and improve team workflows.
  • Alignment with GitLab's values.

Compensation

  • Annual salary range: 126400–314400.

Role Levels

  • Intermediate: Contributes independently to reliability, automation, and operational efficiency within a scoped area.
  • Senior: Drives reliability improvements across multiple projects or services and leads investigations and incident response.
  • Staff: Shapes reliability strategy across teams and defines reusable patterns.
  • Senior Staff: Sets technical direction across a sub-department and drives complex systems initiatives.

Skills

Kubernetes, Terraform, Infrastructure As Code, Go, Ruby, GCP, AWS, CI/CD, GitOps, Observability, Metrics, Logging, SLOs, Incident Response, Automation

Nango

Nango

United States
Staff Engineer, Platform & Infrastructure
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.

Nango

Nango

United States
Staff Platform Engineer
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.

VGS

VGS

United States
Staff Infrastructure Engineer
$145k+/yrRemote8+ YOEDevOps / SRE

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting mission-critical payment systems. Requires 8+ years of distributed-systems experience and deep expertise in infrastructure as code, Kubernetes, automation, and cloud networking.

Grafana Labs

Grafana Labs

United Kingdom
Staff Software Engineer - Databases SRE
£104k+/yrRemote8+ YOEDevOps / SRE

Leads production reliability for Grafana Cloud’s multi-tenant database products, partnering with product engineering teams to improve SLOs, scalability, observability, automation, and incident response. Requires 8+ years of engineering experience, including substantial SRE or production engineering work, plus strong Kubernetes and cloud expertise.

Mozilla

Mozilla

Canada

Senior Staff Performance Engineer, Firefox
CA$149k+/yrRemote7+ YOEDevOps / SRE

Leads Firefox performance engineering by writing code, profiling bottlenecks, improving benchmarks, and guiding cross-functional teams. Requires 7+ years of experience, strong C++ and JavaScript skills, and expertise in performance-critical software, profiling, concurrency, and systems analysis.