Senior Staff Lead Site Reliability Engineer
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
About the job
Responsibilities
- Define and implement service-level indicators (SLIs), service-level objectives (SLOs), and other reliability measures.
- Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services.
- Lead technical responses to complex incidents and drive root-cause analysis through resolution.
- Identify recurring failure modes and work with engineering teams to eliminate them.
- Improve system resilience through automation, testing, capacity planning, and failure recovery.
- Develop tooling and automation to reduce manual operational work.
- Partner with product and platform teams to incorporate reliability requirements into system design.
- Establish incident response practices that improve detection, diagnosis, communication, and recovery.
- Mentor product engineers and SRE and Cloud Engineering teammates.
- Define and manage short- and long-term SRE roadmaps and distribute work across teammates.
Requirements
- 7+ years of experience in SRE, software engineering, infrastructure engineering, or a related field.
- Experience operating production services with defined availability and reliability requirements.
- Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices.
- Experience designing and operating infrastructure in AWS or another major cloud environment.
- Experience with infrastructure as code and automated infrastructure provisioning.
- Experience supporting containerized applications and distributed systems.
- Experience developing operational tooling or automation using Python, Go, or a similar language.
- Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services.
- Experience leading incident response and root-cause analysis across engineering teams.
- Experience leading and executing a team’s technical vision over multi-quarter timelines.
Nice-to-haves
- Experience establishing or maturing an SRE function within an engineering organization.
- Experience with Kubernetes and cloud-native observability systems.
- Experience operating systems in regulated or compliance-driven environments.
- Experience with capacity planning, performance analysis, and cloud cost management.
- Experience supporting shared infrastructure across multiple products or engineering organizations.
- Experience building roadmaps in ticketing systems.
- Experience mentoring other engineers.
Skills
AWS, Infrastructure As Code, Kubernetes, Python, Go, Monitoring, Alerting, Logging, Distributed Systems, Incident Response, Root-Cause Analysis, Capacity Planning, Observability
Similar jobs
DevOps / SRE jobsBuild and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.
Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
Staff DevSecOps Engineer designing and automating security controls across AWS infrastructure, containers, CI/CD, and platform services. Requires 7+ years of related experience plus expertise in cloud security, infrastructure as code, hardened images, vulnerability scanning, identity, and secrets management.