Skip to content
OktaOktaSan Francisco, CA

Staff TDI Site Reliability Engineer, Okta Federal

Staff SRE on Okta's TDI team building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring for national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI with polygraph clearance.

174k – 239k/yr
Hybrid7+ YOEDevOps / SRE

About the role

What you’ll be doing

  • Operate and maintain enterprise grade solutions within air-gapped environments.
  • Build, run, and monitor development tools, pipelines, and infrastructure with a security-first mindset.
  • Operate autonomously within secure facilities.
  • Maintain SLOs/SLIs for workloads with no dependency on external monitoring or SaaS tooling.
  • Own runbooks and incident response procedures tailored to limited external escalation paths.
  • Participate in POA&M remediation and support annual/recurring Authority to Operate activities.
  • Support and run mission critical services depended on by product teams.
  • Deliver excellent internal customer service and advocate for SRE and DevOps practices across teams.
  • Build and operate CI/CD pipelines that function without internet connectivity.

What you’ll bring to the role

  • 7+ years of experience as an SRE, DevOps Engineer, Cloud Automation Engineer, or Systems Engineer with a track record of delivering complex infrastructure projects at scale.
  • Experience with container orchestration and runtime environments, including EKS, ECS Fargate, and general container usage.
  • Proficient in infrastructure automation using Terraform and developing automation tools with Python, while leveraging secure software development practices.
  • Experience with monitoring tools, especially Splunk, CloudWatch, and the Grafana stack.
  • Experience with general networking concepts, such as BGP and IPsec management, and has leveraged AWS networking services, including VPCs, TGWs, and VPC endpoints.
  • Security Clearance: Active U.S. TS/SCI with polygraph.

Additional requirements

  • The selected candidate may be subject to drug testing to the extent required by U.S. Government contracts.

Extra credit

  • Knowledgeable in Linux system administration.
  • Experience in secure and compliant environments (e.g., FedRAMP), with understanding of FIPS, STIG, and data boundary implementations.
  • Working experience operating tools and services in air gapped environments.

Skills

SREDevOpsTerraformPythonEKSECSfargateSplunkCloudWatchGrafanaaws networkingvpctgwBGPipsec

Similar roles

DevOps / SRE jobs
Crusoe

Staff Network Engineer, Deployment

CrusoeDenver, CO

Leads physical and logical deployment of network infrastructure in data centers for AI/HPC, including rack/stack, testing, automation with Python/Ansible, and partner coordination. Requires 8+ years experience with Arista, Juniper, NVIDIA hardware, BGP/EVPN, and physical layer expertise.

174k – 211k/yr
On-site8+ YOEDevOps / SRE
Okta

Staff Site Reliability Engineer - Kubernetes

OktaBellevue, WA +4

Staff SRE builds and manages scalable Kubernetes platforms on AWS, focusing on reliability, automation, cost optimization, and high availability using tools like Helm, Karpenter, and Istio. Requires 5+ years AWS, 4+ years Kubernetes/Helm/Terraform experience.

174k – 267k/yr
Hybrid5+ YOEDevOps / SRE
Grafana Labs

Staff Software Engineer

Grafana LabsUnited States

Staff Platform SysEng building and scaling the Internal Engineering Platform (Kubernetes clusters, infrastructure, tools) that powers Grafana Cloud services (Mimir, Loki, Tempo). Requires proven large-scale distributed systems leadership, cloud-native expertise, reliability ownership, and Go/Python coding.

175k – 210k/yr
Remote7+ YOEDevOps / SRE
Sage

Senior/Staff Site Reliability Engineer

SageNew York, NY

Leads design, operation, and evolution of highly reliable, scalable production infrastructure including cloud, databases, and observability. Drives incident response, SRE practices, automation, and capacity planning for large-scale distributed systems. Requires 7-12+ years in SRE/infrastructure engineering.

175k – 230k/yr
Hybrid7+ YOEDevOps / SRE
Fireworks AI

Member of Technical Staff, Performance Optimization

Fireworks AISan Mateo, CA

Optimizes performance of high-scale systems by analyzing latency, throughput, and resource usage. Requires expertise in profiling, systems programming, and distributed scaling techniques.

175k – 220k/yr
On-siteDevOps / SRE