Skip to content
MongoDBMongoDB

Senior Site Reliability Engineer

The Senior Site Reliability Engineer will operate and improve Kubernetes-based distributed infrastructure for AI application workloads, focusing on scalability, observability, reliability, and tenant isolation. The role requires 6+ years of distributed-systems experience, production Kubernetes expertise, cloud infrastructure knowledge, and strong programming skills.

About the job

Responsibilities

  • Operate and improve multi-tenant Kubernetes infrastructure running customer workloads.
  • Build reliable, resilient, fault-tolerant, available, and self-healing services and infrastructure.
  • Identify and configure metrics to detect incidents and quantify service health, availability, and performance.
  • Participate in a 24/7 on-call rotation to resolve platform infrastructure issues.
  • Mentor early-career SREs and contribute to the team’s operational practices.

Requirements

  • Strong background in software development and operating distributed systems.
  • 6+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language.
  • Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues.
  • Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure.
  • Strong understanding of Linux operating system internals and networking concepts such as TCP/IP, DNS, TLS, and routing.
  • Customer-focused mindset and strong verbal and written technical communication skills.
  • Bias toward efficient processes, operational simplicity, and automation over manual work.
  • Eagerness to learn and a strong technical background.

Nice-to-haves

  • Kubernetes networking experience with technologies such as Istio or Cilium.
  • Production experience with service mesh or edge load balancing.
  • Experience with secure multi-tenant runtime environments at scale.
  • Multi-cloud infrastructure management experience.
  • Experience with virtualization or workload isolation technologies.

Skills

Kubernetes, Python, Go, AWS, GCP, Microsoft Azure, Linux, TCP/IP, DNS, Tls, Routing, Istio, Cilium, Service Mesh, Edge Load Balancing

Okta

Okta

Bengaluru, India

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

GitLab

GitLab

Bengaluru, India

Senior Release Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.

GitLab

GitLab

Bengaluru, India

Senior Site Reliability Engineer - Monitoring and Anomaly Detection
No salary listedRemote5+ YOEDevOps / SRE

Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.