# Staff Site Reliability Engineer - Ecosystem

**Company:** [Okta](https://hotfix.jobs/companies/okta)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 8+ years
**Skills:** AWS, GCP, Kubernetes, Amazon Eks, Google Gke, Terraform, Ansible, Python, Go, Argo Cd, Gitlab Ci, Prometheus, Grafana, Linux, Hashicorp Vault
**Posted:** 2026-08-05

> Leads the design, operation, and modernization of multi-cloud infrastructure across AWS and Google Cloud, with deep Kubernetes, automation, observability, and SRE expertise. The role requires 8+ years in SRE, DevOps, or infrastructure engineering and experience leading large-scale migrations and cross-team reliability initiatives.

## Job Description

## Responsibilities

- Design, build, and operate highly scalable, reliable, and secure infrastructure powering production systems across AWS and Google Cloud.
- Lead reliability and modernization initiatives, including container platform migrations such as ECS to EKS/GKE and microservice enablement across multi-cloud environments.
- Serve as a technical authority in Kubernetes, cloud infrastructure, and modern CI/CD practices including GitOps and automation pipelines.
- Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability.
- Implement and manage infrastructure as code with Terraform and Ansible to automate provisioning, scaling, and configuration management across cloud providers.
- Improve observability, performance, and cost efficiency through monitoring, logging, and alerting systems spanning AWS and Google Cloud.
- Champion SRE practices by defining SLOs and SLIs, conducting blameless postmortems, and improving incident response.
- Lead complex technical projects from conception to completion, managing timelines and dependencies across teams.
- Mentor engineers and foster a culture of reliability, automation, and continuous learning.
- Collaborate with security and compliance partners to ensure infrastructure follows applicable standards, including IAM Federation and Workload Identity.
- Participate in the on-call rotation and use incidents to improve systems and processes.

## Requirements

- 8+ years of experience in SRE, DevOps, or infrastructure engineering roles.
- 3–5 years of production experience with Kubernetes, including EKS, GKE, and related tools such as Helm and Karpenter.
- 3–5 years of experience with AWS and Google Cloud.
- 3–5 years of experience using Terraform to manage multi-cloud infrastructure.
- 3+ years of coding experience in Python, Go, or similar languages.
- Experience leading high-impact migration projects, particularly ECS-to-EKS/GKE migrations, and enabling microservice architectures.
- Experience implementing SLOs/SLIs, performing root-cause analyses, and improving operational resilience.
- Strong Linux and security fundamentals.
- Bachelor’s degree in Computer Science or equivalent hands-on experience.
- Strong communication and problem-solving skills, including experience leading cross-team projects and mentoring peers.

## Technical Expertise

- Cloud-native distributed systems across AWS and Google Cloud.
- Infrastructure as code using Terraform, Ansible, or CloudFormation.
- CI/CD with Argo CD, GitLab CI, or Spinnaker.
- Linux systems and networking fundamentals, including Direct Connect/Interconnect, DNS, routing, and load balancing.
- Databases and caching systems such as RDS/Cloud SQL, Redis/Memorystore, PostgreSQL, and MySQL.
- Observability tools including Prometheus, Grafana, ELK, Loki, OpenTelemetry, and Google Cloud Operations.
- Container security and secrets management using HashiCorp Vault, AWS Secrets Manager, or Google Secret Manager.

## Nice to Have

- Experience in SaaS or high-scale, cloud-native environments.

## Similar jobs

- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/b1b1bdb1-d3f0-47cd-b94b-7781be7399fc) - Together AI - Remote
- [Staff DevSecOps Engineer, Enterprise Technology](https://hotfix.jobs/jobs/85d2a18a-f6f6-4858-a074-0d6b16869dd1) - Okta - Bengaluru, India
- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/cf3c46c0-63dd-4681-b117-16200b400d89) - Together AI - London, United Kingdom
- [Staff Software Engineer](https://hotfix.jobs/jobs/17a92e38-a4f6-4c6b-a6cc-e3bf61bd9955) - Okta - Bengaluru, India
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/d9d546d9-f857-4149-96ca-bf88eceb178d
**Canonical:** https://hotfix.jobs/jobs/d9d546d9-f857-4149-96ca-bf88eceb178d