Staff TDI Site Reliability Engineer, Okta Federal
Staff SRE building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring in isolated environments to support national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI clearance.
About the job
What you’ll be doing
- Operate and maintain enterprise grade solutions within air-gapped environments.
- Build, run, and monitor development tools, pipelines, and infrastructure with a security-first mindset.
- Operate autonomously within secure facilities.
- Maintain SLOs/SLIs for workloads with no dependency on external monitoring or SaaS tooling.
- Own runbooks and incident response procedures tailored to limited external escalation paths.
- Participate in POA&M remediation and support annual/recurring Authority to Operate activities.
- Support and run mission critical services depended on by product teams.
- Deliver excellent internal customer service and advocate for SRE and DevOps practices across teams.
- Build and operate CI/CD pipelines that function without internet connectivity.
What you’ll bring to the role
- 7+ years of experience as an SRE, DevOps Engineer, Cloud Automation Engineer, or Systems Engineer with a track record of delivering complex infrastructure projects at scale.
- Experience with container orchestration and runtime environments, including EKS, ECS Fargate, and general container usage.
- Proficient in infrastructure automation using Terraform and developing automation tools with Python, while leveraging secure software development practices.
- Experience with monitoring tools, especially Splunk, CloudWatch, and the Grafana stack.
- Experience with general networking concepts, such as BGP and IPsec management, and has leveraged AWS networking services, including VPCs, TGWs, and VPC endpoints.
- Security Clearance: Active U.S. TS/SCI with polygraph.
- The selected candidate may be subject to drug testing to the extent required by U.S. Government contracts.
Additional requirements
- U.S. soil status - the employee must be on U.S. soil, which means the 50 states, the District of Columbia, or outlying areas of the United States, as defined in Federal Acquisition Regulation (FAR) 2.101.
- U.S. Security Clearance status - the employee must be able to obtain and maintain a U.S. security clearance (Secret or Top Secret) to the extent required by U.S. Government contracts.
Skills
SRE, DevOps, Terraform, Python, EKS, ECS, Splunk, CloudWatch, Grafana, Aws Networking, BGP, Ipsec, CI/CD
Similar jobs
DevOps / SRE jobsBuild and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.
Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.
Leads reliability engineering for highly available, FedRAMP-compliant cloud services, including infrastructure architecture, automation, observability, incident response, and operational standards. Requires extensive Kubernetes, cloud, software engineering, and cross-team technical leadership experience, plus US-person eligibility and residence on US soil.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.