# Senior Site Reliability Engineer

**Company:** [MongoDB](https://hotfix.jobs/companies/mongodb)
**Location:** Gurugram, India
**Role:** DevOps / SRE
**Experience:** 6+ years
**Skills:** Kubernetes, Python, Go, AWS, GCP, Microsoft Azure, Linux, TCP/IP, DNS, Tls, Routing, Istio, Cilium, Service Mesh, Edge Load Balancing
**Posted:** 2026-08-12

> The Senior Site Reliability Engineer will operate and improve Kubernetes-based distributed infrastructure for AI application workloads, focusing on scalability, observability, reliability, and tenant isolation. The role requires 6+ years of distributed-systems experience, production Kubernetes expertise, cloud infrastructure knowledge, and strong programming skills.

## Job Description

## Responsibilities
- Operate and improve multi-tenant Kubernetes infrastructure running customer workloads.
- Build reliable, resilient, fault-tolerant, available, and self-healing services and infrastructure.
- Identify and configure metrics to detect incidents and quantify service health, availability, and performance.
- Participate in a 24/7 on-call rotation to resolve platform infrastructure issues.
- Mentor early-career SREs and contribute to the team’s operational practices.

## Requirements
- Strong background in software development and operating distributed systems.
- 6+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language.
- Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues.
- Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure.
- Strong understanding of Linux operating system internals and networking concepts such as TCP/IP, DNS, TLS, and routing.
- Customer-focused mindset and strong verbal and written technical communication skills.
- Bias toward efficient processes, operational simplicity, and automation over manual work.
- Eagerness to learn and a strong technical background.

## Nice-to-haves
- Kubernetes networking experience with technologies such as Istio or Cilium.
- Production experience with service mesh or edge load balancing.
- Experience with secure multi-tenant runtime environments at scale.
- Multi-cloud infrastructure management experience.
- Experience with virtualization or workload isolation technologies.

## Similar jobs

- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/0496c897-3cf6-419f-bea7-91e629b4307c) - Okta - Bengaluru, India
- [Senior Release Engineer](https://hotfix.jobs/jobs/a187d17c-c638-43f1-8376-209fe54f2503) - GitLab - Remote
- [Senior Site Reliability Engineer - Monitoring and Anomaly Detection](https://hotfix.jobs/jobs/4d131f28-278c-4b63-a411-189242898dcc) - GitLab - Remote
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior DevOps Engineer](https://hotfix.jobs/jobs/a90d1d14-3f9d-40d5-ae01-365e6600fcab) - ZoomInfo - Bengaluru, India

**Apply:** https://hotfix.jobs/jobs/4d816599-0024-4edb-b23d-a3f80dc7ed04
**Canonical:** https://hotfix.jobs/jobs/4d816599-0024-4edb-b23d-a3f80dc7ed04