Skip to content
OktaOktaBellevue, WA

Staff Site Reliability Engineer

Build and maintain highly available cloud infrastructure edge and CI/CD platforms for Okta's identity platform. Requires 8+ years of cloud operations and Linux systems experience, strong networking knowledge, and proficiency with automation tools like Terraform.

174k – 267k/yr
Hybrid8+ YOEDevOps / SRE

About the role

What you’ll be doing

  • Maintain a highly available cloud infrastructure edge for the Okta identity platform
  • Automate AWS infrastructure with Terraform and/or Chef
  • Evolve the system by introducing changes to improve efficiency, scalability, and velocity

What you’ll bring to the role

  • 8+ years of operations experience configuring, deploying, monitoring and troubleshooting applications and infrastructure in the cloud
  • 8+ years administering or operating within a Linux environment, strong experience using Linux based tooling, and ability to debug systems level problems
  • In-depth understanding of TCP/IP, HTTP, Load Balancing, DNS and other networking protocols
  • Solid understanding and experience with Apache httpd, nginx, Apache Tomcat, or similar
  • Strong problem solving and debugging skills coupled with a desire to take on ownership and responsibility
  • Proficiency in Bash, Python, Golang, or similar. Experienced with git
  • Experience working with Terraform, Ansible, Chef, Puppet or similar automation tools
  • Excellent written and verbal communication skills
  • Willingness to work on-call

Nice-to-haves

  • Experience working in a security-oriented cloud environment
  • Experience debugging software using gdb, strace, ltrace, tcpdump, Wireshark, etc.
  • Experience working with Docker and Kubernetes
  • Experience with PKI / certificate management

Skills

TerraformchefAWSLinuxTCP/IPhttpDNSload balancingapachenginxtomcatBashPythonGoGit

Similar roles

DevOps / SRE jobs
Okta

Staff TDI Site Reliability Engineer, Okta Federal

OktaWashington, DC

Staff SRE building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring in isolated environments to support national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI clearance.

174k – 239k/yr
Hybrid7+ YOEDevOps / SRE
Okta

Staff TDI Site Reliability Engineer, Okta Federal

OktaSan Francisco, CA +1

Staff SRE on Okta's TDI team building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring for national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI with polygraph clearance.

174k – 239k/yr
Hybrid7+ YOEDevOps / SRE
Okta

Staff Site Reliability Engineer - Kubernetes

OktaBellevue, WA +4

Staff SRE builds and manages scalable Kubernetes platforms on AWS, focusing on reliability, automation, cost optimization, and high availability using tools like Helm, Karpenter, and Istio. Requires 5+ years AWS, 4+ years Kubernetes/Helm/Terraform experience.

174k – 267k/yr
Hybrid5+ YOEDevOps / SRE
Grafana Labs

Staff Software Engineer

Grafana LabsUnited States

Staff Platform SysEng building and scaling the Internal Engineering Platform (Kubernetes clusters, infrastructure, tools) that powers Grafana Cloud services (Mimir, Loki, Tempo). Requires proven large-scale distributed systems leadership, cloud-native expertise, reliability ownership, and Go/Python coding.

175k – 210k/yr
Remote7+ YOEDevOps / SRE
Sage

Senior/Staff Site Reliability Engineer

SageNew York, NY

Leads design, operation, and evolution of highly reliable, scalable production infrastructure including cloud, databases, and observability. Drives incident response, SRE practices, automation, and capacity planning for large-scale distributed systems. Requires 7-12+ years in SRE/infrastructure engineering.

175k – 230k/yr
Hybrid7+ YOEDevOps / SRE