# Senior Site Reliability Engineer

**Company:** [Okta](https://hotfix.jobs/companies/okta)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Kubernetes, AWS, GCP, Terraform, Helm, Go, Python, GitOps, Argo CD, Datadog, Splunk, Postgres, Redis, Opensearch, CI/CD
**Posted:** 2026-09-08

> Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

## Job Description

## Responsibilities

### Reliability & Operations
- Design, build, and operate large-scale cloud infrastructure and production services.
- Participate in an on-call rotation supporting highly available customer-facing systems.
- Lead incident response and post-incident reviews focused on systemic improvements.
- Define, measure, and improve SLIs, SLOs, and error budgets.
- Improve service availability, scalability, performance, resilience, and observability.

### Engineering & Automation
- Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
- Eliminate operational toil through automation, tooling, and platform engineering.
- Improve deployment safety and operational workflows through CI/CD and GitOps practices.
- Build self-service platforms, operational guardrails, and developer automation.

### Technical Leadership & Innovation
- Drive reliability initiatives and guide engineers in operational best practices.
- Mentor engineers through design reviews, incident analysis, and knowledge sharing.
- Support architecture and operational decisions with data-driven recommendations.
- Execute projects from conception through production rollout and long-term ownership.
- Apply AI-assisted engineering techniques to operational efficiency, incident response, troubleshooting, and automation.

## Requirements
- Strong experience operating large-scale production services in AWS and/or GCP.
- Deep production Kubernetes expertise, including networking, storage, scheduling, scaling, and workload lifecycle troubleshooting.
- Extensive experience with Infrastructure as Code, including Terraform and Helm.
- Strong software engineering skills in Go and/or Python.
- Experience building automation and internal engineering platforms.
- Experience operating distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
- Understanding of cloud networking, DNS, load balancing, ingress, TLS, service networking, and traffic management.
- Experience with observability platforms, monitoring strategies, and production telemetry.
- Experience leading incident response and improving operations.
- Deep understanding of SLIs, SLOs, error budgets, and capacity planning.
- Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operations.
- Understanding of cloud security, IAM, secrets management, and secure infrastructure design.
- Strong collaboration and communication skills, including experience in globally distributed organizations.
- Experience mentoring engineers and contributing to complex engineering initiatives.

## Nice-to-Haves
- Experience operating SaaS platforms serving large-scale customer workloads.
- Experience in Kubernetes-based microservices environments.
- Experience supporting globally distributed production environments.
- Experience with GitOps and ArgoCD.
- Experience implementing AI-assisted operational tooling or automation.
- Experience with regulated or security-sensitive environments.

## Similar jobs

- [Senior Release Engineer](https://hotfix.jobs/jobs/a187d17c-c638-43f1-8376-209fe54f2503) - GitLab - Remote
- [Senior Site Reliability Engineer - Monitoring and Anomaly Detection](https://hotfix.jobs/jobs/4d131f28-278c-4b63-a411-189242898dcc) - GitLab - Remote
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior DevOps Engineer](https://hotfix.jobs/jobs/a90d1d14-3f9d-40d5-ae01-365e6600fcab) - ZoomInfo - Bengaluru, India
- [Senior DevOps Engineer](https://hotfix.jobs/jobs/a8156c42-7aca-4b0c-9d3b-7a4163c436a4) - ZoomInfo - Bengaluru, India

**Apply:** https://hotfix.jobs/jobs/0496c897-3cf6-419f-bea7-91e629b4307c
**Canonical:** https://hotfix.jobs/jobs/0496c897-3cf6-419f-bea7-91e629b4307c