Skip to content
OktaOkta

Senior Site Reliability Engineer

Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.

About the job

Responsibilities

Reliability & Operations

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in a global on-call rotation and incident response for highly available, customer-facing systems.
  • Lead post-incident reviews focused on systemic improvements.
  • Define, measure, and improve SLIs, SLOs, error budgets, availability, scalability, performance, and resilience.
  • Maintain FedRAMP compliance, security controls, and continuous audit readiness.
  • Improve observability through metrics, logging, tracing, dashboards, and alerting.

Engineering & Automation

  • Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
  • Eliminate operational toil through automation, tooling, and platform engineering.
  • Improve deployment safety and workflows through CI/CD and GitOps practices.
  • Modernize workloads and build self-service platforms, operational guardrails, and developer tooling.

Collaboration & Technical Contribution

  • Support reliability initiatives spanning multiple engineering teams.
  • Guide engineers in operational best practices and reliability engineering principles.
  • Mentor junior and mid-level engineers, conduct code reviews, and provide operational guidance.
  • Contribute to architecture and operational decisions through data-driven recommendations.
  • Drive projects from conception through production rollout and long-term operational ownership.

Requirements

  • Experience operating large-scale production services in AWS and/or GCP.
  • Deep production experience with Linux and Kubernetes, including networking, storage, scheduling, scaling, and workload lifecycle troubleshooting.
  • Extensive experience with Infrastructure as Code, including Terraform and Helm.
  • Strong software engineering skills in Go and/or Python.
  • Experience building automation and internal engineering platforms.
  • Experience operating distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
  • Strong understanding of cloud networking, DNS, load balancing, ingress, TLS, service networking, and traffic management.
  • Experience with observability platforms, monitoring strategies, and production telemetry.
  • Experience operating customer-facing systems subject to SLAs and leading incident response.
  • Understanding of SLIs, SLOs, error budgets, capacity planning, CI/CD, deployment strategies, and automation-first operations.
  • Understanding of cloud security, IAM, secrets management, and secure infrastructure design.
  • US Person status (US citizen or Green Card holder) to support US FedRAMP projects.
  • Strong collaboration, communication, mentoring, and technical leadership skills.

Preferred Qualifications

  • Experience with FedRAMP, SOC 2, HIPAA, or other compliance standards in regulated or government cloud environments.
  • Experience operating SaaS platforms serving large-scale workloads.
  • Experience with Kubernetes-based microservices and globally distributed production environments.
  • Experience with GitOps and ArgoCD.
  • Experience applying AI-assisted engineering or operational automation.

Technology Stack

  • Kubernetes (EKS/GKE)
  • Terraform
  • Helm
  • Git
  • ArgoCD
  • GitOps
  • Go
  • Python
  • Rust
  • Datadog
  • Splunk
  • Grafana
  • PostgreSQL
  • Redis
  • OpenSearch
  • Snowflake

Compensation

  • San Francisco Bay Area annual base salary: $165,000–$227,000 USD.
  • California outside the San Francisco Bay Area, Colorado, Illinois, New York, and Washington annual base salary: $147,000–$202,000 USD.
  • Compensation may also include equity, bonus, and benefits such as health, dental and vision insurance, 401(k), flexible spending accounts, PTO, and parental leave.

Skills

AWS, GCP, Linux, Kubernetes, Terraform, Helm, Go, Python, Rust, Argo CD, GitOps, Datadog, Splunk, Grafana, Postgres

Okta

Okta

Bellevue, WA
Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.

Okta

Okta

Bellevue, WA

Senior Site Reliability Engineer -
$147k+/yrHybrid5+ YOEDevOps / SRE

The Senior Site Reliability Engineer will build and operate secure, scalable infrastructure and Snowflake data systems, automate deployments and operational processes, and lead incident response. The role requires strong coding, Terraform, Kubernetes, CI/CD, and data-platform experience, plus U.S. Person status.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Gumloop

Gumloop

San Francisco, CA
Senior Infrastructure Engineer
$150k+/yrOn-siteDevOps / SRE

Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.

Axle

Axle

Frederick, MD

IT Operations Technical Lead
$150k+/yrHybrid10+ YOEDevOps / SRE

Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.