Build and maintain highly available cloud infrastructure edge and CI/CD platforms for Okta's identity platform. Requires 8+ years of cloud operations and Linux systems experience, strong networking knowledge, and proficiency with automation tools like Terraform.
174k – 267k/yr
Hybrid8+ YOEDevOps / SRE
About the role
What you’ll be doing
Maintain a highly available cloud infrastructure edge for the Okta identity platform
Automate AWS infrastructure with Terraform and/or Chef
Evolve the system by introducing changes to improve efficiency, scalability, and velocity
What you’ll bring to the role
8+ years of operations experience configuring, deploying, monitoring and troubleshooting applications and infrastructure in the cloud
8+ years administering or operating within a Linux environment, strong experience using Linux based tooling, and ability to debug systems level problems
In-depth understanding of TCP/IP, HTTP, Load Balancing, DNS and other networking protocols
Solid understanding and experience with Apache httpd, nginx, Apache Tomcat, or similar
Strong problem solving and debugging skills coupled with a desire to take on ownership and responsibility
Proficiency in Bash, Python, Golang, or similar. Experienced with git
Experience working with Terraform, Ansible, Chef, Puppet or similar automation tools
Excellent written and verbal communication skills
Willingness to work on-call
Nice-to-haves
Experience working in a security-oriented cloud environment
Experience debugging software using gdb, strace, ltrace, tcpdump, Wireshark, etc.
Staff SRE building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring in isolated environments to support national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI clearance.
174k – 239k/yr
Hybrid7+ YOEDevOps / SRE
Staff TDI Site Reliability Engineer, Okta Federal
OktaSan Francisco, CA +1
Staff SRE on Okta's TDI team building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring for national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI with polygraph clearance.
174k – 239k/yr
Hybrid7+ YOEDevOps / SRE
Staff Site Reliability Engineer - Kubernetes
OktaBellevue, WA +4
Staff SRE builds and manages scalable Kubernetes platforms on AWS, focusing on reliability, automation, cost optimization, and high availability using tools like Helm, Karpenter, and Istio. Requires 5+ years AWS, 4+ years Kubernetes/Helm/Terraform experience.
174k – 267k/yr
Hybrid5+ YOEDevOps / SRE
Staff Software Engineer
Grafana LabsUnited States
Staff Platform SysEng building and scaling the Internal Engineering Platform (Kubernetes clusters, infrastructure, tools) that powers Grafana Cloud services (Mimir, Loki, Tempo). Requires proven large-scale distributed systems leadership, cloud-native expertise, reliability ownership, and Go/Python coding.
175k – 210k/yr
Remote7+ YOEDevOps / SRE
Senior/Staff Site Reliability Engineer
SageNew York, NY
Leads design, operation, and evolution of highly reliable, scalable production infrastructure including cloud, databases, and observability. Drives incident response, SRE practices, automation, and capacity planning for large-scale distributed systems. Requires 7-12+ years in SRE/infrastructure engineering.