# Principal Site Reliability Engineer

**Company:** [Okta](https://hotfix.jobs/companies/okta)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 7+ years
**Skills:** AWS, GCP, Kubernetes, Terraform, Helm, GitOps, Argo CD, Go, Python, Datadog, Splunk, Postgres, Redis, Opensearch, Distributed Systems
**Posted:** 2026-07-31

> Leads reliability strategy, architecture, and operational excellence for large-scale cloud products, while building automation and internal platforms. The role requires deep Kubernetes, cloud infrastructure, distributed systems, observability, and reliability engineering expertise, plus strong cross-organizational technical leadership.

## Job Description

## Responsibilities

### Reliability Strategy & Architecture
- Define and drive the reliability strategy for critical product and platform services.
- Establish standards for availability, resilience, observability, incident management, and operational readiness.
- Lead architecture reviews for critical services and platform initiatives.
- Align reliability objectives with business priorities and customer expectations.
- Create frameworks, standards, and operational guardrails for safe operation at scale.
- Guide service architecture toward simplicity, scalability, resilience, and operational excellence.
- Drive initiatives that improve platform maturity and long-term sustainability.

### Product & Platform Leadership
- Own reliability architecture and operational excellence for the Spera / Identity Security Posture Management (ISPM) product area.
- Establish reliability objectives and technical roadmaps with engineering leadership.
- Lead scalability, resiliency, and performance initiatives.
- Build self-service operational capabilities that improve developer productivity while strengthening reliability and security.
- Influence technical direction through data-driven recommendations, engineering expertise, and collaborative leadership.
- Support highly available, large-scale cloud environments as part of an on-call rotation.

### Engineering & Automation
- Design, build, and operate large-scale cloud infrastructure and production services.
- Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
- Eliminate operational toil through automation, tooling, and platform engineering.
- Improve deployment safety and operational workflows through GitOps and infrastructure-as-code practices.
- Modernize existing workloads and align them with evolving platform capabilities.
- Lead engineering initiatives from conception through production rollout and long-term operational ownership.

### Technical Leadership
- Mentor Staff and Senior engineers across multiple teams and organizations.
- Lead technical reviews, design reviews, and operational readiness assessments.
- Build engineering consensus across teams with differing priorities and objectives.
- Develop the next generation of technical leaders.
- Drive adoption of reliability engineering best practices across the Emerging Products Group.
- Share patterns, tooling, and operational practices across Workflows, Inbox, PAM, and ISPM teams.
- Influence technical direction through expertise, collaboration, and execution rather than organizational authority.

### AI & Agentic Operations
- Explore and adopt AI-assisted reliability engineering practices.
- Design and champion agentic systems for troubleshooting, incident response, root-cause analysis, and operational decision-making.
- Evaluate emerging AI technologies and identify opportunities to improve reliability engineering workflows.
- Establish safe, effective, and measurable practices for AI in production operations.
- Reduce operational toil and improve engineering productivity through intelligent automation.

## Technical Requirements
- Extensive experience designing and operating large-scale production systems in AWS and/or GCP.
- Deep expertise with Kubernetes in production environments.
- Experience designing reliability strategies for Kubernetes-based platforms.
- Expertise troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle challenges.
- Extensive experience with infrastructure-as-code technologies such as Terraform and Helm.
- Strong software engineering skills in Go and/or Python.
- Experience building internal platforms, developer tooling, and operational automation.
- Deep understanding of distributed systems architecture and cloud-native application design.
- Understanding of cloud networking fundamentals, including DNS, service discovery, ingress, load balancing, TLS, traffic management, and multi-region architectures.
- Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
- Experience establishing observability standards, monitoring strategies, and operational best practices.
- Experience with or strong interest in AI-assisted engineering and operational automation.
- Strong expertise operating customer-facing production systems at scale.
- Deep understanding of reliability engineering principles, including SLIs, SLOs, error budgets, capacity planning, and resilience engineering.
- Experience leading major incident response efforts and driving long-term operational improvements.
- Strong understanding of CI/CD, GitOps, deployment strategies, and automation-first operations.
- Proven success driving reliability transformations and architectural modernization efforts.
- Ability to balance reliability, scalability, security, and delivery needs.

## Similar jobs

- [Principal Operations Engineer, Mechanical](https://hotfix.jobs/jobs/1033a3a6-548c-4699-8e58-e9cb70cf32cd) - Fluidstack - Remote - $150k – $250k/yr
- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/b1b1bdb1-d3f0-47cd-b94b-7781be7399fc) - Together AI - Remote
- [Staff DevSecOps Engineer, Enterprise Technology](https://hotfix.jobs/jobs/85d2a18a-f6f6-4858-a074-0d6b16869dd1) - Okta - Bengaluru, India
- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/cf3c46c0-63dd-4681-b117-16200b400d89) - Together AI - London, United Kingdom
- [Staff Software Engineer](https://hotfix.jobs/jobs/17a92e38-a4f6-4c6b-a6cc-e3bf61bd9955) - Okta - Bengaluru, India

**Apply:** https://hotfix.jobs/jobs/3636068a-d09f-454d-bab0-74664d3d1f3c
**Canonical:** https://hotfix.jobs/jobs/3636068a-d09f-454d-bab0-74664d3d1f3c