Staff Site Reliability Engineer
Leads reliability engineering for highly available, FedRAMP-compliant cloud services, including infrastructure architecture, automation, observability, incident response, and operational standards. Requires extensive Kubernetes, cloud, software engineering, and cross-team technical leadership experience, plus US-person eligibility and residence on US soil.
About the job
Responsibilities
Reliability & Operations
- Design, build, and operate large-scale cloud infrastructure and production services.
- Participate in a global on-call rotation supporting highly available customer-facing systems.
- Lead incident response and drive post-incident reviews focused on systemic improvements.
- Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Partner with engineering teams to improve availability, scalability, performance, and resilience.
- Maintain FedRAMP compliance, security mandates, and continuous audit readiness.
- Improve observability through metrics, logging, tracing, dashboards, and alerting.
Engineering & Automation
- Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
- Eliminate operational toil through automation, tooling, and platform engineering.
- Improve deployment safety and operational workflows through CI/CD and GitOps practices.
- Modernize workloads and align them with evolving platform capabilities.
- Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security.
Technical Leadership
- Lead complex reliability initiatives spanning multiple engineering teams.
- Guide engineers in operational best practices and reliability engineering principles.
- Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing.
- Influence architecture and operational decisions through data-driven recommendations and engineering expertise.
- Drive projects from conception through production rollout and long-term operational ownership.
Innovation
- Apply AI-assisted engineering techniques to improve operational efficiency, incident response, troubleshooting, and automation.
- Identify emerging technologies that reduce toil and improve engineering productivity.
Requirements
- Extensive experience architecting and leading large-scale production services in AWS and/or GCP.
- Deep expertise defining Kubernetes patterns and Linux-based system standards for enterprise production environments.
- Experience designing multi-region, highly available cloud architectures.
- Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.
- Experience evaluating build-versus-buy decisions and setting long-term technical standards.
- Extensive experience with Infrastructure as Code, including Terraform and Helm.
- Strong software engineering skills in Golang and/or Python.
- Experience building automation and internal engineering platforms.
- Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
- Strong understanding of cloud networking, including DNS, load balancing, ingress, TLS, service networking, and traffic management.
- Experience designing observability frameworks and telemetry-driven operational strategies.
- Experience operating customer-facing production systems subject to SLAs.
- Experience leading incident response and operational improvements.
- Deep understanding of SLIs, SLOs, error budgets, and capacity planning.
- Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operations.
- Ability to balance reliability, scalability, security, and engineering velocity.
- Understanding of cloud security, IAM, secrets management, and secure infrastructure design.
- Experience leading complex engineering initiatives across multiple teams.
- Strong collaboration and communication skills.
- Experience working in globally distributed engineering organizations.
- Ability to define engineering standards, raise technical quality, and mentor junior engineers.
- US Person status (US Citizen or Green Card Holder) and residence on US soil, including the 50 states, the District of Columbia, or applicable outlying areas, are required for FedRAMP projects.
Nice-to-haves
- Experience with FedRAMP, SOC 2, HIPAA, or other compliance standards in regulated or government cloud environments.
- Experience operating SaaS platforms serving large-scale customer workloads.
- Experience with Kubernetes-based microservices environments.
- Experience supporting globally distributed production environments.
- Experience with GitOps and ArgoCD.
- Experience implementing AI-assisted operational tooling or automation workflows.
Tech Stack
- Infrastructure and orchestration: Kubernetes, EKS, GKE, Terraform, Helm, Git, ArgoCD, GitOps
- Programming: Golang, Python, Rust
- Observability: Datadog, Splunk, Grafana
- Data stores: PostgreSQL, Redis, OpenSearch, Snowflake
Compensation
- Annual base salary for candidates in the San Francisco Bay Area: $194,000–$267,000 USD.
- Okta also offers equity where applicable, bonus, and benefits.
Skills
Kubernetes, Amazon Web Services, GCP, Terraform, Helm, Go, Python, Rust, Git, Argo CD, GitOps, Datadog, Splunk, Grafana, Postgres
Similar jobs
DevOps / SRE jobsBuild and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.
Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.