Staff Site Reliability Engineer, Kubernetes w/ active TS/SCI
Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.
About the job
Responsibilities
- Design and implement scalable, reliable network solutions with cross-functional teams.
- Maintain highly available cloud infrastructure at the edge of the identity platform.
- Collect and analyze data to identify root causes of network-specific events.
- Automate AWS infrastructure using Terraform.
- Improve efficiency, scalability, and engineering velocity through system changes.
Requirements
- Active U.S. TS/SCI clearance with full-scope polygraph.
- Experience with Federal and DoD compliance frameworks, including FedRAMP and Impact Level 6 (IL6); familiarity with SC2S/C2S.
- 8+ years of experience in cloud network engineering or a related role.
- In-depth knowledge of TCP/IP networking layers 2–7.
- Ability to implement highly available VPC networks and inter-VPC connectivity.
- Working knowledge of stateless and stateful firewalls, DNS, web application firewalls, and cloud load balancing.
- Deep knowledge of AWS Transit Gateway, Site-to-Site VPN, and Direct Connect.
- Experience with Apache httpd, nginx, Apache Tomcat, or similar technologies.
- Ability to troubleshoot with AWS VPC flow logs, CloudWatch metrics, and packet captures.
- Experience with Terraform and Git.
- Ability to collaborate effectively with multiple stakeholders.
Nice-to-haves
- Bash, Python, or Golang proficiency.
- Palo Alto next-generation virtual firewalls, including security policies, routing, and GlobalProtect.
- Debugging with gdb, strace, ltrace, tcpdump, or Wireshark.
- Docker and Kubernetes experience.
Compensation and benefits
- Annual base salary of $174,000–$238,000 USD.
- Equity, bonus, health, dental and vision insurance, 401(k), flexible spending account, and paid leave, including PTO and parental leave.
Skills
AWS, Terraform, TCP/IP, Vpc, Aws Transit Gateway, Aws Direct Connect, Site-To-Site Vpn, DNS, Load Balancing, Firewalls, FedRAMP, Il6, Kubernetes, Docker, Python
Similar jobs
DevOps / SRE jobsBuild and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.
Leads reliability engineering for highly available, FedRAMP-compliant cloud services, including infrastructure architecture, automation, observability, incident response, and operational standards. Requires extensive Kubernetes, cloud, software engineering, and cross-team technical leadership experience, plus US-person eligibility and residence on US soil.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.