Director, Site Reliability Engineering
Leads India-based Site Reliability Engineering teams responsible for Okta’s cloud platform, databases, networking, Kubernetes, CI/CD, observability, FinOps, and automation. Requires 16+ years in infrastructure or SRE, substantial people-management experience, and expertise in AWS, Kubernetes, Terraform, and reliable SaaS operations.
About the job
Responsibilities
- Build and lead a high-caliber India-based SRE organization supporting Okta’s production fleet.
- Partner with global engineering, product, and infrastructure leaders to deliver resilient, scalable, and secure services.
- Define and execute the India SRE strategy in alignment with global reliability goals.
- Lead post-incident reviews, root-cause analysis, incident management, on-call rotations, and blameless RCAs.
- Implement automation and observability to reduce manual toil and improve operational efficiency.
- Drive adoption of infrastructure as code, container orchestration, and AI within infrastructure operations.
- Hire, mentor, and develop SRE talent across India while building a culture centered on reliability and innovation.
- Foster collaboration across U.S. and EMEA teams and promote continuous learning and process improvement.
- Manage service and business expectations and prioritize resource allocation.
- Develop robust platforms, tooling, and self-service capabilities to accelerate SRE and product engineering.
- Improve cloud infrastructure SDLC processes, including CI/CD, change management, and release management.
Requirements
- 16+ years of experience in site reliability, infrastructure, or production engineering.
- 8+ years of technical leadership and people management experience, including managing managers.
- Experience building or scaling offshore SRE teams that partner with global counterparts.
- 4+ years running an SRE organization supporting a SaaS or cloud service in a public cloud, preferably AWS.
- Strong expertise in automation, observability, performance optimization, and incident response.
- Expertise in cloud-native architectures, Kubernetes, Terraform, and CI/CD pipelines.
- Demonstrated ability to lead cross-functional teams and manage large-scale programs.
- Excellent verbal, written, communication, interpersonal, and cross-cultural collaboration skills.
- Computer Science or related degree, or equivalent experience.
Nice-to-haves
- Experience supporting a multi-cloud environment.
- Experience with AI applied to infrastructure operations.
Compensation and Benefits
- In-person onboarding experience.
- Opportunities for developing talent, fostering connection and community, and supporting social impact.
Skills
AWS, Kubernetes, Terraform, CI/CD, Infrastructure As Code, Observability, Automation, Incident Response, Performance Optimization, Cloud-Native Architecture, SaaS, Multi-Cloud, Root-Cause Analysis, Containerization, Finops
Similar jobs
Engineering Management jobsLeads and builds Starburst’s AI agent platform in India, owning its roadmap, architecture, production operations, and evolution from internal infrastructure to enterprise-facing products. Requires 10+ years of engineering experience, team leadership, platform development, and production LLM systems expertise.
Leads the engineering organization responsible for scaling GitLab.com through cell-based architecture, customer migrations, routing, and multi-cloud infrastructure. The role requires director-level people leadership, distributed-systems expertise, large-scale database and migration experience, and strong asynchronous communication.
Leads a team of Partner Solution Architects across the APAC partner ecosystem, driving strategic relationships, technical alignment, joint solutions, and business growth. Requires extensive partner-management or solution-architecture experience, substantial leadership experience, and a bachelor’s degree or equivalent.
Leads the organization and strategy for Databricks’ large-scale data infrastructure, including billing correctness, disaster recovery, reliability, deployment automation, and data integration. The role requires extensive distributed-systems experience, infrastructure leadership, and experience managing managers.
Leads engineering organization and developer-platform initiatives, including scalable cloud services, AI-assisted development, technical strategy, hiring, and leadership development. Requires 15+ years building distributed systems and experience managing high-performance engineering teams.