Site Reliability Engineer, Infrastructure Platforms
Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.
About the job
Responsibilities
- Keep user-facing services and production systems reliable, scalable, and efficient.
- Build automation and tooling to reduce toil and replace manual work with repeatable, infrastructure-as-code-driven workflows.
- Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling.
- Write and maintain infrastructure as code, shipping changes safely through CI/CD and GitOps.
- Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately.
- Contribute to observability using metrics, logs, and SLOs to detect symptoms early.
- Participate in incident response and post-incident reviews, turning learnings into automation and process improvements.
- Document runbooks, architecture decisions, and reviews.
Requirements
- Experience keeping production systems reliable while combining an operations mindset with software engineering practice.
- Experience building new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators or controllers, and production automation or services from scratch.
- Ability to read, debug, and reason about code, including behavior, performance, and failure modes.
- Experience with infrastructure as code, Kubernetes, and its ecosystem.
- Hands-on experience with GCP or AWS.
- Familiarity with observability practices, including metrics, logging, alerting, SLOs, and SLIs.
- Experience with on-call and incident response, using a structured troubleshooting approach.
- Strong written communication and ability to work independently in an asynchronous, distributed environment.
- Track record of using automation and AI to reduce toil and improve team practices.
Skills
Kubernetes, Terraform, Infrastructure As Code, Go, Ruby, GCP, AWS, CI/CD, GitOps, Observability, SLOs, Slis, Incident Response, On-Call, Automation
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.
Build and operate cloud-agnostic deployment and observability infrastructure across public clouds and on-premises environments. The role requires 5+ years of infrastructure experience, strong networking and IaC expertise, and ownership of production systems and cross-functional projects.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.