# Staff Site Reliability Engineer

**Company:** [Okta](https://hotfix.jobs/companies/okta)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 7+ years
**Skills:** Kubernetes, Amazon Web Services, GCP, Terraform, Helm, Go, Python, GitOps, Argo Cd, Datadog, Splunk, Postgres, Redis, Opensearch, CI/CD
**Posted:** 2026-08-05

> Leads reliability engineering for large-scale, customer-facing cloud services, driving automation, observability, incident response, and operational excellence. The role requires deep Kubernetes, cloud infrastructure, infrastructure-as-code, software engineering, and distributed systems expertise, along with cross-team technical leadership and mentoring.

## Job Description

## Responsibilities

### Reliability & Operations
- Design, build, and operate large-scale cloud infrastructure and production services.
- Participate in an on-call rotation supporting highly available customer-facing systems.
- Lead incident response efforts and drive post-incident reviews focused on systemic improvements.
- Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Partner with engineering teams to improve service availability, scalability, performance, and resilience.
- Improve observability through metrics, logging, tracing, dashboards, and alerting.

### Engineering & Automation
- Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
- Eliminate operational toil through automation, tooling, and platform engineering.
- Improve deployment safety and operational workflows through CI/CD and GitOps practices.
- Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities.
- Build self-service platforms, operational guardrails, and automation that improve developer velocity while maintaining reliability and security.

### Technical Leadership
- Lead complex reliability initiatives spanning multiple engineering teams.
- Guide engineers in adopting operational best practices and reliability engineering principles.
- Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing.
- Influence architecture and operational decisions through data-driven recommendations and engineering expertise.
- Drive projects from conception through production rollout and long-term operational ownership.

### Innovation
- Explore and apply AI-assisted engineering techniques to improve operational efficiency, incident response, troubleshooting, and automation.
- Identify opportunities to leverage emerging technologies to reduce toil and improve engineering productivity.

## Requirements

- Strong experience operating large-scale production services in AWS and/or GCP.
- Deep expertise with Kubernetes in production environments.
- Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.
- Extensive experience with infrastructure-as-code technologies such as Terraform and Helm.
- Strong software engineering skills in Go and/or Python.
- Experience building automation and internal engineering platforms.
- Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
- Strong understanding of cloud networking fundamentals, including DNS, load balancing, ingress, TLS, service networking, and traffic management.
- Experience with observability platforms, monitoring strategies, and production telemetry.
- Experience with or strong interest in AI-assisted engineering and operational automation.
- Experience operating customer-facing production systems.
- Experience leading incident response and driving operational improvements.
- Deep understanding of reliability engineering concepts, including SLIs, SLOs, error budgets, and capacity planning.
- Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operational practices.
- Ability to balance reliability, scalability, security, and engineering velocity.
- Understanding of cloud security fundamentals, IAM, secrets management, and secure infrastructure design.
- Demonstrated success leading complex engineering initiatives across multiple teams.
- Strong collaboration and communication skills.
- Experience working effectively within globally distributed engineering organizations spanning multiple time zones and cultures.
- Experience mentoring engineers and elevating technical capabilities within an organization.
- Ability to influence technical direction through expertise, partnership, and execution.

## Nice-to-Haves

- Experience implementing operational controls and best practices in regulated or security-sensitive environments.
- Experience operating SaaS platforms serving large-scale customer workloads.
- Experience working within Kubernetes-based microservices environments.
- Experience supporting globally distributed production environments.
- Experience with GitOps and Argo CD.
- Experience implementing AI-assisted operational tooling or automation workflows.

## Technology Stack

- **Infrastructure and orchestration:** Kubernetes, Amazon EKS, Google Kubernetes Engine, Terraform, Helm, Git, Argo CD, GitOps
- **Programming:** Go, Python
- **Observability:** Datadog, Splunk
- **Data stores:** PostgreSQL, Redis, OpenSearch

## Similar jobs

- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/b1b1bdb1-d3f0-47cd-b94b-7781be7399fc) - Together AI - Remote
- [Staff DevSecOps Engineer, Enterprise Technology](https://hotfix.jobs/jobs/85d2a18a-f6f6-4858-a074-0d6b16869dd1) - Okta - Bengaluru, India
- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/cf3c46c0-63dd-4681-b117-16200b400d89) - Together AI - London, United Kingdom
- [Staff Software Engineer](https://hotfix.jobs/jobs/17a92e38-a4f6-4c6b-a6cc-e3bf61bd9955) - Okta - Bengaluru, India
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/7867ac6b-2e65-4abb-bfc1-11b36d517cb9
**Canonical:** https://hotfix.jobs/jobs/7867ac6b-2e65-4abb-bfc1-11b36d517cb9