# Senior Staff Lead Site Reliability Engineer

**Company:** [Shield AI](https://hotfix.jobs/companies/shield-ai)
**Location:** San Diego, CA
**Role:** DevOps / SRE
**Salary:** $190k – $280k/yr
**Experience:** 7+ years
**Skills:** AWS, Infrastructure As Code, Kubernetes, Python, Go, Monitoring, Alerting, Logging, Distributed Systems, Incident Response, Root-Cause Analysis, Capacity Planning, Observability
**Posted:** 2026-09-03

> Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

## Job Description

## Responsibilities
- Define and implement service-level indicators (SLIs), service-level objectives (SLOs), and other reliability measures.
- Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services.
- Lead technical responses to complex incidents and drive root-cause analysis through resolution.
- Identify recurring failure modes and work with engineering teams to eliminate them.
- Improve system resilience through automation, testing, capacity planning, and failure recovery.
- Develop tooling and automation to reduce manual operational work.
- Partner with product and platform teams to incorporate reliability requirements into system design.
- Establish incident response practices that improve detection, diagnosis, communication, and recovery.
- Mentor product engineers and SRE and Cloud Engineering teammates.
- Define and manage short- and long-term SRE roadmaps and distribute work across teammates.

## Requirements
- 7+ years of experience in SRE, software engineering, infrastructure engineering, or a related field.
- Experience operating production services with defined availability and reliability requirements.
- Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices.
- Experience designing and operating infrastructure in AWS or another major cloud environment.
- Experience with infrastructure as code and automated infrastructure provisioning.
- Experience supporting containerized applications and distributed systems.
- Experience developing operational tooling or automation using Python, Go, or a similar language.
- Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services.
- Experience leading incident response and root-cause analysis across engineering teams.
- Experience leading and executing a team’s technical vision over multi-quarter timelines.

## Nice-to-haves
- Experience establishing or maturing an SRE function within an engineering organization.
- Experience with Kubernetes and cloud-native observability systems.
- Experience operating systems in regulated or compliance-driven environments.
- Experience with capacity planning, performance analysis, and cloud cost management.
- Experience supporting shared infrastructure across multiple products or engineering organizations.
- Experience building roadmaps in ticketing systems.
- Experience mentoring other engineers.

## Similar jobs

- [Staff Platform Engineer](https://hotfix.jobs/jobs/abd59d24-c218-4099-b1f0-d9427e499a31) - OpenSea - Remote - $190k – $345k/yr
- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/63668c94-fd42-45fb-9699-5d39d126e387) - Komodo Health - Remote - $187k – $265k/yr
- [Senior Staff Infrastructure Engineer](https://hotfix.jobs/jobs/f99a2e1a-7c43-4a02-ba1c-9d76880c0acd) - VGS - Remote - $185k – $290k/yr
- [Staff Network Engineer, Operations](https://hotfix.jobs/jobs/19d77267-e1da-45b5-9ed9-fce44a90d734) - Crusoe - San Francisco, CA - $195k – $235k/yr
- [Staff DevSecOps Engineer](https://hotfix.jobs/jobs/593e5319-afa3-4408-bf94-a4885cd3c100) - Shield AI - San Mateo, CA - $182k – $274k/yr

**Apply:** https://hotfix.jobs/jobs/c5516e9e-9df1-44ef-b77a-caf084a94fc8
**Canonical:** https://hotfix.jobs/jobs/c5516e9e-9df1-44ef-b77a-caf084a94fc8