Skip to content
OktaOkta

Staff Site Reliability Engineer

Leads reliability engineering for highly available, FedRAMP-compliant cloud services, including infrastructure architecture, automation, observability, incident response, and operational standards. Requires extensive Kubernetes, cloud, software engineering, and cross-team technical leadership experience, plus US-person eligibility and residence on US soil.

About the job

Responsibilities

Reliability & Operations

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in a global on-call rotation supporting highly available customer-facing systems.
  • Lead incident response and drive post-incident reviews focused on systemic improvements.
  • Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Partner with engineering teams to improve availability, scalability, performance, and resilience.
  • Maintain FedRAMP compliance, security mandates, and continuous audit readiness.
  • Improve observability through metrics, logging, tracing, dashboards, and alerting.

Engineering & Automation

  • Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
  • Eliminate operational toil through automation, tooling, and platform engineering.
  • Improve deployment safety and operational workflows through CI/CD and GitOps practices.
  • Modernize workloads and align them with evolving platform capabilities.
  • Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security.

Technical Leadership

  • Lead complex reliability initiatives spanning multiple engineering teams.
  • Guide engineers in operational best practices and reliability engineering principles.
  • Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing.
  • Influence architecture and operational decisions through data-driven recommendations and engineering expertise.
  • Drive projects from conception through production rollout and long-term operational ownership.

Innovation

  • Apply AI-assisted engineering techniques to improve operational efficiency, incident response, troubleshooting, and automation.
  • Identify emerging technologies that reduce toil and improve engineering productivity.

Requirements

  • Extensive experience architecting and leading large-scale production services in AWS and/or GCP.
  • Deep expertise defining Kubernetes patterns and Linux-based system standards for enterprise production environments.
  • Experience designing multi-region, highly available cloud architectures.
  • Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.
  • Experience evaluating build-versus-buy decisions and setting long-term technical standards.
  • Extensive experience with Infrastructure as Code, including Terraform and Helm.
  • Strong software engineering skills in Golang and/or Python.
  • Experience building automation and internal engineering platforms.
  • Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
  • Strong understanding of cloud networking, including DNS, load balancing, ingress, TLS, service networking, and traffic management.
  • Experience designing observability frameworks and telemetry-driven operational strategies.
  • Experience operating customer-facing production systems subject to SLAs.
  • Experience leading incident response and operational improvements.
  • Deep understanding of SLIs, SLOs, error budgets, and capacity planning.
  • Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operations.
  • Ability to balance reliability, scalability, security, and engineering velocity.
  • Understanding of cloud security, IAM, secrets management, and secure infrastructure design.
  • Experience leading complex engineering initiatives across multiple teams.
  • Strong collaboration and communication skills.
  • Experience working in globally distributed engineering organizations.
  • Ability to define engineering standards, raise technical quality, and mentor junior engineers.
  • US Person status (US Citizen or Green Card Holder) and residence on US soil, including the 50 states, the District of Columbia, or applicable outlying areas, are required for FedRAMP projects.

Nice-to-haves

  • Experience with FedRAMP, SOC 2, HIPAA, or other compliance standards in regulated or government cloud environments.
  • Experience operating SaaS platforms serving large-scale customer workloads.
  • Experience with Kubernetes-based microservices environments.
  • Experience supporting globally distributed production environments.
  • Experience with GitOps and ArgoCD.
  • Experience implementing AI-assisted operational tooling or automation workflows.

Tech Stack

  • Infrastructure and orchestration: Kubernetes, EKS, GKE, Terraform, Helm, Git, ArgoCD, GitOps
  • Programming: Golang, Python, Rust
  • Observability: Datadog, Splunk, Grafana
  • Data stores: PostgreSQL, Redis, OpenSearch, Snowflake

Compensation

  • Annual base salary for candidates in the San Francisco Bay Area: $194,000–$267,000 USD.
  • Okta also offers equity where applicable, bonus, and benefits.

Skills

Kubernetes, Amazon Web Services, GCP, Terraform, Helm, Go, Python, Rust, Git, Argo CD, GitOps, Datadog, Splunk, Grafana, Postgres

Okta

Okta

Bellevue, WA
Staff Site Reliability Engineer - Kubernetes
$174k+/yrHybrid7+ YOEDevOps / SRE

Build and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.

Okta

Okta

Maryland
Staff Site Reliability Engineer, Kubernetes w/ active TS/SCI
$174k+/yrHybrid8+ YOEDevOps / SRE

Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Attentive

Attentive

United States

Staff Site Reliability Engineer
$180k+/yrRemote7+ YOEDevOps / SRE

Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.

Shield AI

Shield AI

United States

Sr. Staff Platform/Data Reliability Engineer, Databricks
$180k+/yrRemote12+ YOEDevOps / SRE

Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.