Manager, Site Reliability Engineering
Lead and grow teams of SREs and managers responsible for Edge networking, Kubernetes platform, CI/CD, observability, and automation tooling that powers Okta's high-availability IDaaS platform on AWS. Drive DevOps maturity, self-service capabilities, reliability, and infrastructure-as-code practices.
About the job
What you’ll be doing
- Managing a team of SRE’s supporting various workloads and teams that support our IDaaS platform.
- Drive the microservice journey, DevOps maturity, and workload reliability in tandem with architects and teams across the organization.
- Accelerate the velocity of SRE and product engineering by developing powerful tooling, intuitive self-service capabilities, and robust self-healing patterns.
- Lead, mentor, and grow a high-performing team of engineers and managers across platform, infrastructure, and shared services domains.
- Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints.
- Improve SDLC processes for Cloud infrastructure as a code, including the maturity of CI/CD pipelines, change and release management.
- Manage service and business expectations and prioritize resource allocation.
- Maintain a deep knowledge of industry best practices, evolving trends, and technologies.
What you’ll bring to the role
- 3+ years of experience in technical leadership & people management.
- Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared service at scale.
- Experience running large-scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS. Experience supporting a multi-Cloud environment will be a plus.
- Strong expertise in cloud-native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines.
- Strong background and hands-on experience in SW development, PaaS and automation.
- Deep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM etc.) in a large scale environment.
- Effective verbal, written communication and interpersonal skills.
- Computer Science Degree or related degree or equivalent experience.
Skills
Kubernetes, Terraform, AWS, CI/CD, Observability, Grafana, Splunk, DevOps, Agile, Iac, Paas, Microservices
Similar jobs
Engineering Management jobsLeads a platform engineering team responsible for shared systems, identity, permissions, session management, and messaging infrastructure. The manager owns delivery and architecture, develops engineers, embeds AI into development workflows, and partners closely with Product and Design.
Engineering manager leading a data platform team that allocates cloud costs to products and customers. The role requires experience with reproducible batch pipelines, versioned financial metrics, cross-functional delivery, and cloud cost, billing, revenue, or financial data systems.
Leads a hands-on team building and maintaining third-party security and enterprise integrations. The manager coaches engineers, guides technical decisions and delivery, and contributes to backend systems, APIs, and production troubleshooting.
Leads the engineering team building an AI-powered referral coordination product integrating document processing, voice automation, and EHR systems. The role combines people management, hands-on technical leadership, scalable architecture, customer collaboration, and healthcare interoperability expertise.
Leads and develops a 4–5 person engineering squad while continuing to write production code and owning delivery, architecture, and operational health. The role focuses on safely shipping LLM-powered financial experiences and requires people-management experience, AI-native development fluency, and hands-on software engineering skills.