# Site Reliability Engineer , Storage Layer Services

**Company:** [MongoDB](https://hotfix.jobs/companies/mongodb)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 6+ years
**Skills:** Python, Go, Kubernetes, AWS, GCP, Microsoft Azure, Linux, TCP/IP, DNS, Tls, Routing, Distributed Systems, Database Systems, Containerization
**Posted:** 2026-08-12

> This SRE will operate and improve MongoDB Atlas’s multi-tenant distributed storage infrastructure, focusing on reliability, performance, observability, automation, and incident response. The role requires 6+ years of distributed-systems experience plus expertise in storage or databases, Kubernetes, cloud platforms, Linux, and networking.

## Job Description

## Responsibilities

- Work on multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs.
- Build reliable services and infrastructure that are available, resilient, fault-tolerant, and self-healing.
- Identify and configure key metrics to detect incidents and quantify service health, availability, and performance.
- Participate in a 24/7 on-call rotation to resolve issues involving storage infrastructure.
- Become an expert in infrastructure performance, optimizing from the application level through the kernel.

## Requirements

- 6+ years of experience working on software development and operating distributed systems.
- Proficiency in Python, Go, or a similar language.
- Experience operating or supporting stateful storage or database systems at scale, including durability, consistency, and recovery trade-offs.
- Customer-focused mindset and an emphasis on operational efficiency.
- Preference for automation over manual processes and reducing operational toil through software solutions.
- Experience using and extending containerization technologies, particularly Kubernetes.
- Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform, or Azure.
- Understanding of Linux operating system internals and networking concepts such as TCP/IP, DNS, TLS, and routing.

## Nice-to-haves

- Experience leading major architectural shifts, including migrations from legacy storage stacks to new multi-tenant storage architectures.
- Experience planning and executing large-scale data and workload migrations with strict availability and durability requirements.
- Experience managing and scaling infrastructure across AWS, Google Cloud Platform, or Azure.
- Experience designing secure, multi-tenant runtime environments at scale.

## Similar jobs

- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Software Engineer, DevOps](https://hotfix.jobs/jobs/bfede940-874a-457b-a7d5-afaa6be82cff) - Muck Rack - Remote - €95k – €110k/yr
- [Senior Database Administrator - Core Infrastructure](https://hotfix.jobs/jobs/84ebc795-fb55-4998-94cf-dd630cc51e8a) - Kraken - Remote
- [Senior Production Engineer](https://hotfix.jobs/jobs/9ee5879e-954d-4681-ad0c-816d7151f874) - Lightspark - Remote - $200k – $238k/yr
- [Senior Network Production Operations Engineer](https://hotfix.jobs/jobs/51d2d612-1ae9-4982-a371-05e810e9dc67) - Crusoe - Dublin, Ireland

**Apply:** https://hotfix.jobs/jobs/2ee927f2-5b05-4326-87c9-3acb08a64764
**Canonical:** https://hotfix.jobs/jobs/2ee927f2-5b05-4326-87c9-3acb08a64764