Site Reliability Engineer , Storage Layer Services
This SRE will operate and improve MongoDB Atlas’s multi-tenant distributed storage infrastructure, focusing on reliability, performance, observability, automation, and incident response. The role requires 6+ years of distributed-systems experience plus expertise in storage or databases, Kubernetes, cloud platforms, Linux, and networking.
About the job
Responsibilities
- Work on multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs.
- Build reliable services and infrastructure that are available, resilient, fault-tolerant, and self-healing.
- Identify and configure key metrics to detect incidents and quantify service health, availability, and performance.
- Participate in a 24/7 on-call rotation to resolve issues involving storage infrastructure.
- Become an expert in infrastructure performance, optimizing from the application level through the kernel.
Requirements
- 6+ years of experience working on software development and operating distributed systems.
- Proficiency in Python, Go, or a similar language.
- Experience operating or supporting stateful storage or database systems at scale, including durability, consistency, and recovery trade-offs.
- Customer-focused mindset and an emphasis on operational efficiency.
- Preference for automation over manual processes and reducing operational toil through software solutions.
- Experience using and extending containerization technologies, particularly Kubernetes.
- Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform, or Azure.
- Understanding of Linux operating system internals and networking concepts such as TCP/IP, DNS, TLS, and routing.
Nice-to-haves
- Experience leading major architectural shifts, including migrations from legacy storage stacks to new multi-tenant storage architectures.
- Experience planning and executing large-scale data and workload migrations with strict availability and durability requirements.
- Experience managing and scaling infrastructure across AWS, Google Cloud Platform, or Azure.
- Experience designing secure, multi-tenant runtime environments at scale.
Skills
Python, Go, Kubernetes, AWS, GCP, Microsoft Azure, Linux, TCP/IP, DNS, Tls, Routing, Distributed Systems, Database Systems, Containerization
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.
Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Operates and scales Crusoe Cloud’s global edge, backbone, and data center networks supporting GPU-based HPC workloads. The role requires extensive production networking experience, strong protocol and observability expertise, automation skills, and participation in 24/7 on-call support.