# Staff Site Reliability Engineer

**Company:** [Okta](https://hotfix.jobs/companies/okta)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 8+ years
**Skills:** AWS, GCP, Kubernetes, Amazon Eks, Google Gke, Terraform, Ansible, Python, Go, CI/CD, Argo Cd, Linux, Redis, Prometheus, Grafana
**Posted:** 2026-07-31

> Leads the design, operation, and modernization of multi-cloud infrastructure across AWS and Google Cloud. The role requires deep Kubernetes expertise, SRE practices, infrastructure as code, automation, observability, and at least eight years of relevant experience.

## Job Description

## Responsibilities
- Design, build, and operate highly scalable, reliable, and secure infrastructure powering production systems across AWS and Google Cloud.
- Lead reliability and modernization initiatives, including container platform migrations such as ECS to EKS/GKE and microservice enablement across multi-cloud environments.
- Serve as a technical authority in Kubernetes, cloud infrastructure, and modern CI/CD practices including GitOps and automation pipelines.
- Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability.
- Implement and manage infrastructure as code with Terraform and Ansible to automate provisioning, scaling, and configuration management across cloud providers.
- Improve observability, performance, and cost efficiency through monitoring, logging, and alerting across AWS and Google Cloud.
- Define SLOs and SLIs, conduct blameless postmortems, and continuously improve incident response.
- Lead complex technical projects from conception to completion, managing timelines and technical dependencies across teams.
- Mentor engineers and foster a culture of reliability, automation, and continuous learning.
- Collaborate with security and compliance partners on infrastructure standards, including IAM Federation and Workload Identity.
- Participate in the on-call rotation and use incidents to improve systems and processes.

## Requirements
- 8+ years in SRE, DevOps, or infrastructure engineering roles.
- 3–5 years of production experience with Kubernetes, including EKS and GKE, and ecosystem tools such as Helm and Karpenter.
- 3–5 years of experience with AWS and Google Cloud.
- 3–5 years using Terraform to manage multi-cloud infrastructure.
- 5+ years of coding experience in Python, Go, or similar languages.
- Hands-on experience architecting and operating cloud-native distributed systems.
- Experience leading ECS-to-EKS/GKE migration projects and enabling microservice architectures.
- Proficiency with Terraform, Ansible, or CloudFormation.
- Advanced understanding of CI/CD pipelines, Linux systems, networking fundamentals, and Redis.
- Experience managing databases and caching systems such as RDS, Cloud SQL, Redis/Memorystore, PostgreSQL, and MySQL.
- Experience with observability tools including Prometheus, Grafana, ELK, Loki, OpenTelemetry, and Google Cloud Operations.
- Working knowledge of container security, secrets management, and production compliance.
- Strong communication and problem-solving skills, including cross-team project leadership and mentoring.
- Strong Linux and security fundamentals.
- Bachelor’s degree in Computer Science or equivalent hands-on experience.

## Nice to Have
- Experience in SaaS or high-scale, cloud-native environments.

## Benefits
- In-person onboarding experience.
- Well-being support.
- Social impact opportunities.
- Talent development and community-building programs.

## Similar jobs

- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/b1b1bdb1-d3f0-47cd-b94b-7781be7399fc) - Together AI - Remote
- [Staff DevSecOps Engineer, Enterprise Technology](https://hotfix.jobs/jobs/85d2a18a-f6f6-4858-a074-0d6b16869dd1) - Okta - Bengaluru, India
- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/cf3c46c0-63dd-4681-b117-16200b400d89) - Together AI - London, United Kingdom
- [Staff Software Engineer](https://hotfix.jobs/jobs/17a92e38-a4f6-4c6b-a6cc-e3bf61bd9955) - Okta - Bengaluru, India
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/9165ace9-7e75-49b7-95a3-ade2d9da054a
**Canonical:** https://hotfix.jobs/jobs/9165ace9-7e75-49b7-95a3-ade2d9da054a