# Staff TDI Site Reliability Engineer, Okta Federal

**Company:** [Okta](https://hotfix.jobs/companies/okta)
**Location:** San Francisco, CA, Washington, DC
**Role:** DevOps / SRE
**Salary:** $174k – $239k/yr
**Experience:** 7+ years
**Skills:** SRE, DevOps, Terraform, Python, EKS, ECS, fargate, Splunk, CloudWatch, Grafana, aws networking, vpc, tgw, BGP, ipsec
**Posted:** 2026-07-17

> Staff SRE on Okta's TDI team building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring for national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI with polygraph clearance.

## Job Description

## What you’ll be doing

- Operate and maintain enterprise grade solutions within air-gapped environments.
- Build, run, and monitor development tools, pipelines, and infrastructure with a security-first mindset.
- Operate autonomously within secure facilities.
- Maintain SLOs/SLIs for workloads with no dependency on external monitoring or SaaS tooling.
- Own runbooks and incident response procedures tailored to limited external escalation paths.
- Participate in POA&M remediation and support annual/recurring Authority to Operate activities.
- Support and run mission critical services depended on by product teams.
- Deliver excellent internal customer service and advocate for SRE and DevOps practices across teams.
- Build and operate CI/CD pipelines that function without internet connectivity.

## What you’ll bring to the role

- 7+ years of experience as an SRE, DevOps Engineer, Cloud Automation Engineer, or Systems Engineer with a track record of delivering complex infrastructure projects at scale.
- Experience with container orchestration and runtime environments, including EKS, ECS Fargate, and general container usage.
- Proficient in infrastructure automation using Terraform and developing automation tools with Python, while leveraging secure software development practices.
- Experience with monitoring tools, especially Splunk, CloudWatch, and the Grafana stack.
- Experience with general networking concepts, such as BGP and IPsec management, and has leveraged AWS networking services, including VPCs, TGWs, and VPC endpoints.
- Security Clearance: Active U.S. TS/SCI with polygraph.

## Additional requirements

- The selected candidate may be subject to drug testing to the extent required by U.S. Government contracts.

## Extra credit

- Knowledgeable in Linux system administration.
- Experience in secure and compliant environments (e.g., FedRAMP), with understanding of FIPS, STIG, and data boundary implementations.
- Working experience operating tools and services in air gapped environments.

## Similar roles

- [Staff TDI Site Reliability Engineer, Okta Federal](https://hotfix.jobs/jobs/279c1ebe-3a0e-4d0a-bdf0-d67dbcabde3b) - Okta - Washington, DC - $174k – $239k/yr
- [Staff Site Reliability Engineer - Kubernetes](https://hotfix.jobs/jobs/ff325b64-e4d7-47b2-a05e-4da96444e68c) - Okta - Bellevue, WA - $174k – $267k/yr
- [Staff Software Engineer](https://hotfix.jobs/jobs/3a484da9-09cb-408e-a132-e6f664620a16) - Grafana Labs - Remote - $175k – $210k/yr
- [Senior/Staff Site Reliability Engineer](https://hotfix.jobs/jobs/176c9f9e-3a92-449a-b4cb-39cf88da222b) - Sage - New York, NY - $175k – $230k/yr
- [Member of Technical Staff, AI Training Infrastructure](https://hotfix.jobs/jobs/3b60d4a2-35a8-4163-8505-a2811ac7ceae) - Fireworks AI - San Mateo, CA - $175k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/b690f9a0-0815-41a5-8755-2d8806f8ec4b
**Canonical:** https://hotfix.jobs/jobs/b690f9a0-0815-41a5-8755-2d8806f8ec4b