Skip to content
OktaOkta

Senior Site Reliability Engineer

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

About the job

Responsibilities

Reliability & Operations

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in an on-call rotation supporting highly available customer-facing systems.
  • Lead incident response and post-incident reviews focused on systemic improvements.
  • Define, measure, and improve SLIs, SLOs, and error budgets.
  • Improve service availability, scalability, performance, resilience, and observability.

Engineering & Automation

  • Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
  • Eliminate operational toil through automation, tooling, and platform engineering.
  • Improve deployment safety and operational workflows through CI/CD and GitOps practices.
  • Build self-service platforms, operational guardrails, and developer automation.

Technical Leadership & Innovation

  • Drive reliability initiatives and guide engineers in operational best practices.
  • Mentor engineers through design reviews, incident analysis, and knowledge sharing.
  • Support architecture and operational decisions with data-driven recommendations.
  • Execute projects from conception through production rollout and long-term ownership.
  • Apply AI-assisted engineering techniques to operational efficiency, incident response, troubleshooting, and automation.

Requirements

  • Strong experience operating large-scale production services in AWS and/or GCP.
  • Deep production Kubernetes expertise, including networking, storage, scheduling, scaling, and workload lifecycle troubleshooting.
  • Extensive experience with Infrastructure as Code, including Terraform and Helm.
  • Strong software engineering skills in Go and/or Python.
  • Experience building automation and internal engineering platforms.
  • Experience operating distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
  • Understanding of cloud networking, DNS, load balancing, ingress, TLS, service networking, and traffic management.
  • Experience with observability platforms, monitoring strategies, and production telemetry.
  • Experience leading incident response and improving operations.
  • Deep understanding of SLIs, SLOs, error budgets, and capacity planning.
  • Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operations.
  • Understanding of cloud security, IAM, secrets management, and secure infrastructure design.
  • Strong collaboration and communication skills, including experience in globally distributed organizations.
  • Experience mentoring engineers and contributing to complex engineering initiatives.

Nice-to-Haves

  • Experience operating SaaS platforms serving large-scale customer workloads.
  • Experience in Kubernetes-based microservices environments.
  • Experience supporting globally distributed production environments.
  • Experience with GitOps and ArgoCD.
  • Experience implementing AI-assisted operational tooling or automation.
  • Experience with regulated or security-sensitive environments.

Skills

Kubernetes, AWS, GCP, Terraform, Helm, Go, Python, GitOps, Argo CD, Datadog, Splunk, Postgres, Redis, Opensearch, CI/CD

GitLab

GitLab

Bengaluru, India

Senior Release Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.

GitLab

GitLab

Bengaluru, India

Senior Site Reliability Engineer - Monitoring and Anomaly Detection
No salary listedRemote5+ YOEDevOps / SRE

Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Leads the design, automation, and reliability of large-scale, multi-cloud infrastructure supporting search, NoSQL, and AI-driven workloads. Requires 7+ years in infrastructure, DevOps, or SRE, plus deep Kubernetes, Terraform, Linux, and distributed data-systems expertise.