Senior Manager, Site Reliability Engineering (Federal)
Lead and mentor multiple SRE teams overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation tooling for Okta’s high-scale SaaS infrastructure on AWS.
About the job
What you'll be doing
- Managing a team of SRE’s supporting our various workloads operating in private sector environments.
- Drive the microservice journey, DevOps maturity, and workload reliability in tandem with architects and teams across the organization.
- Accelerate the velocity of SRE and product engineering by developing powerful tooling, intuitive self-service capabilities, and robust self-healing patterns.
- Lead, mentor, and grow a high-performing team of engineers and managers across platform, infrastructure, and shared services domains.
- Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints.
- Improve SDLC processes for Cloud infrastructure as a code, including the maturity of CI/CD pipelines, change and release management.
- Manage service and business expectations and prioritize resource allocation.
- Maintain a deep knowledge of industry best practices, evolving trends, and technologies.
What you’ll bring to the role
- 3+ years of experience in technical leadership & people management.
- Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared service at scale.
- Experience running large-scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS. Experience supporting a multi-Cloud environment will be a plus.
- Strong expertise in cloud-native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines.
- Strong background and hands-on experience in SW development, PaaS and automation.
- Deep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM etc.) in a large scale environment.
- Effective verbal, written communication and interpersonal skills.
- Computer Science Degree or related degree or equivalent experience.
Additional requirements
- This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.
Skills
Kubernetes, Terraform, AWS, CI/CD, Grafana, Splunk, DevOps, Observability, Infrastructure As Code, Agile
Similar jobs
DevOps / SRE jobsBuild and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.