# Staff Site Reliability Engineer

**Company:** [MongoDB](https://hotfix.jobs/companies/mongodb)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 10+ years
**Skills:** Kubernetes, Python, Go, Distributed Systems, AWS, GCP, Microsoft Azure, Multi-Cloud, Incident Response, SLOs, Capacity Planning, Observability, Networking, Containers, Virtual Machines
**Posted:** 2026-08-12

> Provides technical leadership for the reliability architecture and operational foundations of a multi-region, multi-cloud platform for AI applications. The role requires deep Kubernetes and distributed-systems expertise, infrastructure programming, cloud knowledge, and experience setting SRE standards and mentoring engineers.

## Job Description

## Responsibilities
- Own the reliability architecture of the platform across regions and cloud providers.
- Collaborate with platform teams, providing internal support and guidance on operability, capacity, and best practices.
- Set operational standards for on-call quality, incident response, and SLO discipline.
- Mentor and technically develop the SRE team.
- Participate in a 24/7 on-call rotation to resolve platform infrastructure issues.

## Requirements
- 10+ years of experience working on software and operating distributed systems.
- Deep Kubernetes expertise, including designing or evolving multi-cluster platforms.
- Proficiency in Python, Go, or a similar programming language.
- Understanding of workload isolation at the systems level, including containers, virtual machines, and trade-offs for running untrusted code.
- Customer-focused mindset.
- Preference for automation over manual processes and strong operational efficiency.
- Intimate familiarity with infrastructure primitives of at least one of AWS, Google Cloud, or Microsoft Azure, with the ability to reason about differences between them.
- Track record of driving infrastructure architecture across teams and mentoring engineers.

## Similar jobs

- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/b1b1bdb1-d3f0-47cd-b94b-7781be7399fc) - Together AI - Remote
- [Staff DevSecOps Engineer, Enterprise Technology](https://hotfix.jobs/jobs/85d2a18a-f6f6-4858-a074-0d6b16869dd1) - Okta - Bengaluru, India
- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/cf3c46c0-63dd-4681-b117-16200b400d89) - Together AI - London, United Kingdom
- [Staff Software Engineer](https://hotfix.jobs/jobs/17a92e38-a4f6-4c6b-a6cc-e3bf61bd9955) - Okta - Bengaluru, India
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/2890d60d-314e-4a5f-905f-07fc7bfbed16
**Canonical:** https://hotfix.jobs/jobs/2890d60d-314e-4a5f-905f-07fc7bfbed16