# Senior Site Reliability Engineer

**Company:** [Okta](https://hotfix.jobs/companies/okta)
**Location:** San Francisco, CA
**Role:** DevOps / SRE
**Salary:** $147k – $227k/yr
**Experience:** 5+ years
**Skills:** AWS, GCP, Linux, Kubernetes, Terraform, Helm, Go, Python, Rust, Argo CD, GitOps, Datadog, Splunk, Grafana, Postgres
**Posted:** 2026-08-27

> Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.

## Job Description

## Responsibilities

### Reliability & Operations
- Design, build, and operate large-scale cloud infrastructure and production services.
- Participate in a global on-call rotation and incident response for highly available, customer-facing systems.
- Lead post-incident reviews focused on systemic improvements.
- Define, measure, and improve SLIs, SLOs, error budgets, availability, scalability, performance, and resilience.
- Maintain FedRAMP compliance, security controls, and continuous audit readiness.
- Improve observability through metrics, logging, tracing, dashboards, and alerting.

### Engineering & Automation
- Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
- Eliminate operational toil through automation, tooling, and platform engineering.
- Improve deployment safety and workflows through CI/CD and GitOps practices.
- Modernize workloads and build self-service platforms, operational guardrails, and developer tooling.

### Collaboration & Technical Contribution
- Support reliability initiatives spanning multiple engineering teams.
- Guide engineers in operational best practices and reliability engineering principles.
- Mentor junior and mid-level engineers, conduct code reviews, and provide operational guidance.
- Contribute to architecture and operational decisions through data-driven recommendations.
- Drive projects from conception through production rollout and long-term operational ownership.

### Requirements
- Experience operating large-scale production services in AWS and/or GCP.
- Deep production experience with Linux and Kubernetes, including networking, storage, scheduling, scaling, and workload lifecycle troubleshooting.
- Extensive experience with Infrastructure as Code, including Terraform and Helm.
- Strong software engineering skills in Go and/or Python.
- Experience building automation and internal engineering platforms.
- Experience operating distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
- Strong understanding of cloud networking, DNS, load balancing, ingress, TLS, service networking, and traffic management.
- Experience with observability platforms, monitoring strategies, and production telemetry.
- Experience operating customer-facing systems subject to SLAs and leading incident response.
- Understanding of SLIs, SLOs, error budgets, capacity planning, CI/CD, deployment strategies, and automation-first operations.
- Understanding of cloud security, IAM, secrets management, and secure infrastructure design.
- US Person status (US citizen or Green Card holder) to support US FedRAMP projects.
- Strong collaboration, communication, mentoring, and technical leadership skills.

### Preferred Qualifications
- Experience with FedRAMP, SOC 2, HIPAA, or other compliance standards in regulated or government cloud environments.
- Experience operating SaaS platforms serving large-scale workloads.
- Experience with Kubernetes-based microservices and globally distributed production environments.
- Experience with GitOps and ArgoCD.
- Experience applying AI-assisted engineering or operational automation.

### Technology Stack
- Kubernetes (EKS/GKE)
- Terraform
- Helm
- Git
- ArgoCD
- GitOps
- Go
- Python
- Rust
- Datadog
- Splunk
- Grafana
- PostgreSQL
- Redis
- OpenSearch
- Snowflake

### Compensation
- San Francisco Bay Area annual base salary: $165,000–$227,000 USD.
- California outside the San Francisco Bay Area, Colorado, Illinois, New York, and Washington annual base salary: $147,000–$202,000 USD.
- Compensation may also include equity, bonus, and benefits such as health, dental and vision insurance, 401(k), flexible spending accounts, PTO, and parental leave.

## Similar jobs

- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/0a9de6d2-d738-400d-9f4e-a3fc9057c3e4) - Okta - Bellevue, WA - $147k – $202k/yr
- [Senior Site Reliability Engineer -](https://hotfix.jobs/jobs/f306b084-941a-46e3-b0b2-524ac878c929) - Okta - Bellevue, WA - $147k – $202k/yr
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Infrastructure Engineer](https://hotfix.jobs/jobs/8a829a2b-c825-45b8-ac7e-8a92504dfad2) - Gumloop - San Francisco, CA - $150k – $300k/yr
- [IT Operations Technical Lead](https://hotfix.jobs/jobs/d1895f34-2afe-4c92-a77b-3d8e01b815ef) - Axle - Frederick, MD - $150k – $170k/yr

**Apply:** https://hotfix.jobs/jobs/b5f3d205-95e6-4655-8a41-935cc01b0e48
**Canonical:** https://hotfix.jobs/jobs/b5f3d205-95e6-4655-8a41-935cc01b0e48