Skip to content
OktaOkta

Staff Site Reliability Engineer - Kubernetes

Build and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.

About the job

Responsibilities

  • Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms for production workloads.
  • Build, manage, and optimize AWS infrastructure, including EKS, ECS, S3, VPCs, RDS, IAM, and related services.
  • Create, maintain, and manage Helm charts for production Kubernetes deployments.
  • Implement and manage Karpenter for dynamic Kubernetes cluster scaling and resource optimization.
  • Configure and manage Istio for service-to-service communication, security, observability, traffic management, service discovery, and policy enforcement.
  • Automate infrastructure and application deployment, scaling, and management through CI/CD pipelines.
  • Respond to incidents and troubleshoot performance, availability, and security issues.
  • Implement secure cloud infrastructure with appropriate access controls, network security, and compliance practices.
  • Document Kubernetes platform setup, operational procedures, and best practices, and share knowledge across teams.

Requirements

  • 5+ years of experience with AWS.
  • 4+ years of experience with Kubernetes and Helm.
  • 4+ years of experience with Terraform.
  • Experience with multi-region cloud environments and cloud-native architectures.
  • Strong expertise creating and managing Kubernetes platforms, including highly available clusters, networking, and storage.
  • Hands-on experience deploying and managing applications with Helm.
  • Experience with Karpenter for dynamic Kubernetes scaling.
  • Experience managing and securing Istio, including traffic management, security, and observability.
  • Proficiency with CI/CD pipelines and automation tools such as Jenkins, GitLab, CircleCI, Ansible, and Spinnaker.
  • Strong scripting and automation skills in Python, Bash, or Go.
  • Experience with monitoring, logging, and alerting tools such as Prometheus, Grafana, CloudWatch, and ELK Stack.
  • Ability to meet U.S. Person status requirements for access to federal environments or protected federal data.
  • Ability to complete in-person onboarding and travel to the San Francisco, CA headquarters or Chicago office during the first week.

Nice-to-haves

  • Knowledge of cloud and Kubernetes security practices, including RBAC, encryption, and compliance frameworks.
  • Familiarity with Docker and containerization.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
  • CKA, CKAD, or AWS Certified DevOps Engineer certification.

Skills

Kubernetes, Helm, Terraform, AWS, Amazon Eks, Amazon Ecs, Amazon S3, Amazon Vpc, Amazon Rds, IAM, Karpenter, Istio, CI/CD, Python, Bash

Okta

Okta

Maryland
Staff Site Reliability Engineer, Kubernetes w/ active TS/SCI
$174k+/yrHybrid8+ YOEDevOps / SRE

Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.

Okta

Okta

Bellevue, WA
Staff Site Reliability Engineer
$174k+/yrHybrid7+ YOEDevOps / SRE

Leads reliability engineering for highly available, FedRAMP-compliant cloud services, including infrastructure architecture, automation, observability, incident response, and operational standards. Requires extensive Kubernetes, cloud, software engineering, and cross-team technical leadership experience, plus US-person eligibility and residence on US soil.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Attentive

Attentive

United States

Staff Site Reliability Engineer
$180k+/yrRemote7+ YOEDevOps / SRE

Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.

Shield AI

Shield AI

United States

Sr. Staff Platform/Data Reliability Engineer, Databricks
$180k+/yrRemote12+ YOEDevOps / SRE

Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.