Skip to content
OktaOktaWashington, DC

Staff TDI Site Reliability Engineer, Okta Federal

Staff SRE building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring in isolated environments to support national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI clearance.

174k – 239k/yr
Hybrid7+ YOEDevOps / SRE

About the role

What you’ll be doing

  • Operate and maintain enterprise grade solutions within air-gapped environments.
  • Build, run, and monitor development tools, pipelines, and infrastructure with a security-first mindset.
  • Operate autonomously within secure facilities.
  • Maintain SLOs/SLIs for workloads with no dependency on external monitoring or SaaS tooling.
  • Own runbooks and incident response procedures tailored to limited external escalation paths.
  • Participate in POA&M remediation and support annual/recurring Authority to Operate activities.
  • Support and run mission critical services depended on by product teams.
  • Deliver excellent internal customer service and advocate for SRE and DevOps practices across teams.
  • Build and operate CI/CD pipelines that function without internet connectivity.

What you’ll bring to the role

  • 7+ years of experience as an SRE, DevOps Engineer, Cloud Automation Engineer, or Systems Engineer with a track record of delivering complex infrastructure projects at scale.
  • Experience with container orchestration and runtime environments, including EKS, ECS Fargate, and general container usage.
  • Proficient in infrastructure automation using Terraform and developing automation tools with Python, while leveraging secure software development practices.
  • Experience with monitoring tools, especially Splunk, CloudWatch, and the Grafana stack.
  • Experience with general networking concepts, such as BGP and IPsec management, and has leveraged AWS networking services, including VPCs, TGWs, and VPC endpoints.
  • Security Clearance: Active U.S. TS/SCI with polygraph.
  • The selected candidate may be subject to drug testing to the extent required by U.S. Government contracts.

Additional requirements

  • U.S. soil status - the employee must be on U.S. soil, which means the 50 states, the District of Columbia, or outlying areas of the United States, as defined in Federal Acquisition Regulation (FAR) 2.101.
  • U.S. Security Clearance status - the employee must be able to obtain and maintain a U.S. security clearance (Secret or Top Secret) to the extent required by U.S. Government contracts.

Skills

SREDevOpsTerraformPythonEKSECSSplunkCloudWatchGrafanaaws networkingBGPipsecCI/CD

Similar roles

DevOps / SRE jobs
Okta

Staff TDI Site Reliability Engineer, Okta Federal

OktaSan Francisco, CA +1

Staff SRE on Okta's TDI team building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring for national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI with polygraph clearance.

174k – 239k/yr
Hybrid7+ YOEDevOps / SRE
Okta

Staff Site Reliability Engineer - Kubernetes

OktaBellevue, WA +4

Staff SRE builds and manages scalable Kubernetes platforms on AWS, focusing on reliability, automation, cost optimization, and high availability using tools like Helm, Karpenter, and Istio. Requires 5+ years AWS, 4+ years Kubernetes/Helm/Terraform experience.

174k – 267k/yr
Hybrid5+ YOEDevOps / SRE
Grafana Labs

Staff Software Engineer

Grafana LabsUnited States

Staff Platform SysEng building and scaling the Internal Engineering Platform (Kubernetes clusters, infrastructure, tools) that powers Grafana Cloud services (Mimir, Loki, Tempo). Requires proven large-scale distributed systems leadership, cloud-native expertise, reliability ownership, and Go/Python coding.

175k – 210k/yr
Remote7+ YOEDevOps / SRE
Sage

Senior/Staff Site Reliability Engineer

SageNew York, NY

Leads design, operation, and evolution of highly reliable, scalable production infrastructure including cloud, databases, and observability. Drives incident response, SRE practices, automation, and capacity planning for large-scale distributed systems. Requires 7-12+ years in SRE/infrastructure engineering.

175k – 230k/yr
Hybrid7+ YOEDevOps / SRE
Fireworks AI

Member of Technical Staff, AI Training Infrastructure

Fireworks AISan Mateo, CA

Builds and optimizes scalable infrastructure for AI model training on large GPU clusters. Requires expertise in distributed systems, Python/C++, and ML frameworks.

175k – 220k/yr
On-siteDevOps / SRE