Sr. Staff Lead Site Reliability Engineer
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.
About the job
Responsibilities
- Define and implement service-level indicators (SLIs), service-level objectives (SLOs), and other reliability measures.
- Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services.
- Lead technical response to complex incidents and drive root-cause analysis through resolution.
- Identify recurring failure modes and partner with engineering teams to eliminate them.
- Improve system resilience through automation, testing, capacity planning, and failure recovery.
- Develop tooling and automation to reduce manual operational work.
- Incorporate reliability requirements into system design with product and platform teams.
- Establish incident-response practices that improve detection, diagnosis, communication, and recovery.
- Mentor product engineers and SRE and Cloud Engineering teammates.
- Define and manage short- and long-term SRE roadmaps and distribute work across teammates.
Required Qualifications
- 7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields.
- Experience operating production services with defined availability and reliability requirements.
- Experience implementing SLIs, SLOs, monitoring, alerting, and incident-response practices.
- Experience designing and operating infrastructure in AWS or another major cloud environment.
- Experience with infrastructure as code and automated infrastructure provisioning.
- Experience supporting containerized applications and distributed systems.
- Experience developing operational tooling or automation using Python, Go, or a similar language.
- Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services.
- Experience leading incident response and root-cause analysis across engineering teams.
- Experience leading and executing a team's technical vision over multi-quarter timelines.
Preferred Qualifications
- Experience establishing or maturing an SRE function within an engineering organization.
- Experience with Kubernetes and cloud-native observability systems.
- Experience operating systems in regulated or compliance-driven environments.
- Experience with capacity planning, performance analysis, and cloud cost management.
- Experience supporting shared infrastructure across multiple products or engineering organizations.
- Experience building roadmaps in ticketing systems.
- Experience mentoring other engineers.
Skills
AWS, Slis, SLOs, Monitoring, Alerting, Logging, Distributed Systems, Infrastructure As Code, Kubernetes, Python, Go, Incident Response, Root-Cause Analysis, Capacity Planning, Cloud Cost Management
Similar jobs
DevOps / SRE jobsLeads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.
Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.