# Site Reliability Engineer, Infrastructure Platforms

**Company:** [GitLab](https://hotfix.jobs/companies/gitlab)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Kubernetes, Terraform, Infrastructure As Code, Go, Ruby, GCP, AWS, CI/CD, GitOps, Observability, SLOs, Slis, Incident Response, On-Call, Automation
**Posted:** 2026-09-08

> Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.

## Job Description

## Responsibilities
- Keep user-facing services and production systems reliable, scalable, and efficient.
- Build automation and tooling to reduce toil and replace manual work with repeatable, infrastructure-as-code-driven workflows.
- Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling.
- Write and maintain infrastructure as code, shipping changes safely through CI/CD and GitOps.
- Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately.
- Contribute to observability using metrics, logs, and SLOs to detect symptoms early.
- Participate in incident response and post-incident reviews, turning learnings into automation and process improvements.
- Document runbooks, architecture decisions, and reviews.

## Requirements
- Experience keeping production systems reliable while combining an operations mindset with software engineering practice.
- Experience building new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators or controllers, and production automation or services from scratch.
- Ability to read, debug, and reason about code, including behavior, performance, and failure modes.
- Experience with infrastructure as code, Kubernetes, and its ecosystem.
- Hands-on experience with GCP or AWS.
- Familiarity with observability practices, including metrics, logging, alerting, SLOs, and SLIs.
- Experience with on-call and incident response, using a structured troubleshooting approach.
- Strong written communication and ability to work independently in an asynchronous, distributed environment.
- Track record of using automation and AI to reduce toil and improve team practices.

## Similar jobs

- [Software Engineer: Resiliency - Deploy at Scale](https://hotfix.jobs/jobs/21bb6f4f-daea-46ab-bac8-fff083d17451) - Cloudflare - London, United Kingdom
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Member of Technical Staff](https://hotfix.jobs/jobs/d6c912e7-2f16-4a86-8738-980dd0b47cd0) - Perplexity - Remote - $220k – $405k/yr
- [Infrastructure Engineer](https://hotfix.jobs/jobs/e7b74501-38ae-4cd1-8b26-91e042b41460) - Writer - London, United Kingdom
- [Infrastructure Software Engineer, Apps Platform](https://hotfix.jobs/jobs/5c50edd8-ecfb-4db4-86e8-8698ff8827cd) - Scale AI - London, United Kingdom

**Apply:** https://hotfix.jobs/jobs/e28d2ba0-df16-41d6-9c60-e694ee353996
**Canonical:** https://hotfix.jobs/jobs/e28d2ba0-df16-41d6-9c60-e694ee353996