Senior Manager, Site Reliability Engineering - Infrastructure Platform
Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.
About the job
The Infrastructure Platform and Shared Services Team
Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput and 99.999 availability.
As the Sr. Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, CI/CD, Observability, automation platform & tooling.
What you’ll be doing
- Lead the Infra platform and shared services org and various initiatives across SRE & Infrastructure organization.
- Lead the DevOps transformation, microservice journey, and next generation Infra platform capabilities in partnership with architects and product engineering
- Build a world-class observability platform and monitoring capabilities enabled with self-service
- Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and intuitive self-service capabilities.
- Own the design and operation of scalable, self-service Cloud infrastructure platforms (e.g., Kubernetes, service mesh, CI/CD pipelines, IaC & Edge Infrastructure)
- Lead, mentor, and grow a high-performing team of engineers and managers across platform, infrastructure, and shared services domains.
- Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints.
- Improve SDLC processes for Cloud infrastructure as a code, including the maturity of CI/CD pipelines, change and release management
- Manage service and business expectations and prioritize resource allocation
- Maintain a deep knowledge of industry best practices, evolving trends, and technologies
What you’ll bring to the role
- 6+ years of experience in technical leadership & people management
- Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared service at scale
- 3+ years of experience running large-scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS. Experience supporting a multi-Cloud environment will be a plus.
- Strong expertise in cloud-native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines
- Strong background and hands-on experience in SW development, PaaS and automation
- Deep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM etc.) in a large scale environment.
- Demonstrated ability to lead cross-functional teams and manage large-scale programs
- Effective verbal, written communication and interpersonal skills
- Computer Science Degree or related degree or equivalent experience
Skills
Kubernetes, AWS, Terraform, CI/CD, Grafana, Splunk, DevOps, Iac, Observability, Service Mesh
Similar jobs
DevOps / SRE jobsBuild and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.
Leads cross-functional technical initiatives and builds scalable business operations and customer-facing systems. Requires Python, system design, production engineering experience, and strong stakeholder collaboration; platform, AWS, SaaS, and analytics experience are preferred.
Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.