# Manager, Site Reliability Engineering

**Company:** [Okta](https://hotfix.jobs/companies/okta)
**Location:** New York, NY, Washington, DC
**Role:** Engineering Management
**Salary:** $182k – $251k/yr
**Experience:** 8+ years
**Skills:** AWS, Azure, Terraform, Kubernetes, Containers, Microservices, Go, Python, Databases, Observability, Incident Response, Infrastructure Automation
**Posted:** 2026-08-26

> Leads the Site Reliability Engineering team for Auth0, setting technical direction, improving platform resilience, and guiding incident response at scale. Requires 8+ years of industry experience, 3+ years of team leadership, and deep expertise in cloud-native infrastructure, automation, and SRE practices.

## Job Description

## Responsibilities
- Lead the SRE team's technical direction and translate organizational vision into actionable roadmaps.
- Drive complex, cross-functional initiatives across product and platform teams.
- Participate hands-on in 24/7 on-call rotations, troubleshooting and remediating incidents on critical systems.
- Build infrastructure resilience through monitoring, alerting, and automation improvements.
- Establish reliability practices and standards centered on observability, resilience, and software engineering rigor.
- Mentor and develop SRE talent through pair programming, design discussions, and code reviews.
- Represent reliability in architectural reviews and strategic planning.

## Requirements
- 8+ years of total industry experience.
- 3+ years of hands-on team leadership in SRE or software engineering roles.
- Experience in cloud-native environments and architectures, including containers, Kubernetes, microservices, and databases.
- Expertise with AWS, Azure, and Terraform.
- Strong programming skills in Go or Python.
- Experience building and maintaining production-grade tools, automation, and infrastructure solutions.
- Knowledge of SRE principles, blameless incident response, systematic problem-solving, and software engineering approaches to operational challenges.
- Excellent verbal and written communication skills.
- Experience leading high-performing teams in globally distributed, remote-first environments.
- Strategic vision, technical depth, leadership ability, and a commitment to mentoring senior engineers.
- Ability to submit documentation establishing U.S. Person status upon hire.

## Nice-to-haves
- Experience leading reliability initiatives that improved uptime and reduced incident response times at scale.
- Contributions to open-source infrastructure or observability tooling.
- Experience designing incident response programs and runbook automation.

## Compensation
- Annual base salary: **$182,000–$250,800 USD**.
- Equity, bonus, health, dental and vision insurance, 401(k), flexible spending account, PTO, and parental leave may be available according to applicable plans and policies.

## Similar jobs

- [Engineering Manager, Infrastructure](https://hotfix.jobs/jobs/88d32298-9e30-4957-9203-f090c5bc4348) - Formation Bio - New York, NY - $186k – $232k/yr
- [Manager II, Config Deployments](https://hotfix.jobs/jobs/dcac1d6c-8cd1-4831-9484-1f7f04dd8b5d) - Pinterest - Remote - $177k – $365k/yr
- [Senior Manager, Technical Delivery](https://hotfix.jobs/jobs/7deb9822-0b9a-4ed6-ab34-d42c0284fa88) - Snowflake - Chicago, IL - $177k – $232k/yr
- [Tech Lead Manager, Human Agent Tooling](https://hotfix.jobs/jobs/101619fb-3b65-4305-80c7-8bfd4a7e04fc) - Chime - San Francisco, CA - $187k – $259k/yr
- [Senior Manager, Data Center Facility Operations](https://hotfix.jobs/jobs/7e033bdd-2ee7-40d3-8ea5-361888f926c8) - Crusoe - Shakopee, MN - $175k – $190k/yr

**Apply:** https://hotfix.jobs/jobs/c03b9ab9-86b8-489d-8aa6-a56f54e43f1d
**Canonical:** https://hotfix.jobs/jobs/c03b9ab9-86b8-489d-8aa6-a56f54e43f1d