Skip to content
ReltioReltio

Senior Manager Engineering - SRE

Leads a global 24×7 DevOps and Customer Enablement organization supporting highly available SaaS operations. The role requires 12+ years in DevOps, SRE, cloud operations, or production engineering, including substantial experience managing engineering teams and driving reliability, incident response, automation, and customer outcomes.

About the job

Responsibilities

  • Lead 24×7 follow-the-sun operations as the front line for infrastructure alerts and Incident Management, including on-call coverage, rapid response, service restoration, escalation, and cross-region handoffs.
  • Define and improve operational KPIs covering service availability, incident response and RCA, MTTR, SLA compliance, customer issues and escalations, Service Help Desk health, backlog reduction, automation adoption, and operational efficiency.
  • Define and execute Customer Enablement strategy focused on operational excellence, cloud platform reliability, automation, infrastructure readiness, go-live success, and engineering best practices.
  • Lead major customer-impacting incidents and coordinate Engineering, Cloud Platform, Customer Enablement, Support, and Product teams toward resolution and stakeholder communication.
  • Identify recurring customer issues and operational trends, then drive permanent fixes, preventive improvements, and platform automation.
  • Establish governance for Incident Management, Problem Management, Change Management, on-call operations, and RCA processes.
  • Lead customer onboarding, infrastructure readiness, feature enablement, go-live governance, and Cloud Platform Service Help Desk operations.
  • Build cross-functional ownership models, operating mechanisms, service reviews, and escalation paths.
  • Drive Infrastructure as Code, AI-assisted operations, runbook maturity, self-service capabilities, workflow automation, and continuous improvement.
  • Build and lead high-performing global DevOps teams while developing technical leaders and fostering accountability and operational excellence.

Requirements

  • Engineering degree in Computer Science or a related technical field.
  • 12+ years of experience in DevOps, SRE, Cloud Operations, or Production Engineering, including 5+ years leading engineering teams or managers.
  • Experience leading globally distributed 24×7 follow-the-sun DevOps and Customer Enablement operations for highly available SaaS platforms.
  • Hands-on experience with AWS, Google Cloud Platform, or Microsoft Azure.
  • Expertise in Kubernetes, containerized platforms, Terraform, CI/CD, GitOps, and automation.
  • Strong Linux, networking, troubleshooting, distributed systems, and cloud architecture fundamentals.
  • Experience with observability, Incident Management, on-call operations, RCA, SLA/SLO governance, and production reliability.
  • Experience leading Customer Enablement operations, customer onboarding, infrastructure readiness, go-live governance, and operational support for enterprise SaaS customers.
  • Experience defining and improving operational KPIs and translating data into customer and business outcomes.
  • Experience leading major customer-impacting incidents and cross-functional resolution.
  • Experience establishing governance for Incident Management, Problem Management, Change Management, RCA, and continuous service improvement.
  • Ability to influence senior stakeholders, drive cross-functional alignment, and deliver results without direct authority.

Skills

AWS, GCP, Microsoft Azure, Kubernetes, Terraform, Infrastructure As Code, CI/CD, GitOps, Linux, Networking, Distributed Systems, Cloud Architecture, Observability, Incident Management, Rca

Okta

Okta

Bengaluru, India

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

GitLab

GitLab

Bengaluru, India

Senior Release Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.

GitLab

GitLab

Bengaluru, India

Senior Site Reliability Engineer - Monitoring and Anomaly Detection
No salary listedRemote5+ YOEDevOps / SRE

Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.