Skip to content
GitLabGitLab

Site Reliability Engineer, Infrastructure Platforms

Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.

About the job

Responsibilities

  • Keep user-facing services and production systems reliable, scalable, and efficient.
  • Build automation and tooling to reduce toil and replace manual work with repeatable, infrastructure-as-code-driven workflows.
  • Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling.
  • Write and maintain infrastructure as code, shipping changes safely through CI/CD and GitOps.
  • Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately.
  • Contribute to observability using metrics, logs, and SLOs to detect symptoms early.
  • Participate in incident response and post-incident reviews, turning learnings into automation and process improvements.
  • Document runbooks, architecture decisions, and reviews.

Requirements

  • Experience keeping production systems reliable while combining an operations mindset with software engineering practice.
  • Experience building new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators or controllers, and production automation or services from scratch.
  • Ability to read, debug, and reason about code, including behavior, performance, and failure modes.
  • Experience with infrastructure as code, Kubernetes, and its ecosystem.
  • Hands-on experience with GCP or AWS.
  • Familiarity with observability practices, including metrics, logging, alerting, SLOs, and SLIs.
  • Experience with on-call and incident response, using a structured troubleshooting approach.
  • Strong written communication and ability to work independently in an asynchronous, distributed environment.
  • Track record of using automation and AI to reduce toil and improve team practices.

Skills

Kubernetes, Terraform, Infrastructure As Code, Go, Ruby, GCP, AWS, CI/CD, GitOps, Observability, SLOs, Slis, Incident Response, On-Call, Automation

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Writer

Writer

London, United Kingdom

Infrastructure Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.

Scale AI

Scale AI

London, United Kingdom

Infrastructure Software Engineer, Apps Platform
No salary listedOn-site5+ YOEDevOps / SRE

Build and operate cloud-agnostic deployment and observability infrastructure across public clouds and on-premises environments. The role requires 5+ years of infrastructure experience, strong networking and IaC expertise, and ownership of production systems and cross-functional projects.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.