Team Lead, Site Reliability Engineering - Storage Layer Service
Leads a team of SREs for MongoDB's Storage Layer Services, defining SLOs, capacity plans, and roadmaps for multi-tenant distributed storage systems underpinning Atlas. Requires 10+ years in distributed systems and 2+ years managing teams, with expertise in Kubernetes and IaC tools.
About the job
Responsibilities
- Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers
- Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs
- Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges
- Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations
Requirements
- 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams
- Customer-focused mindset, treating internal developers as your primary users
- Value efficiency in processes and operations, and have a track record of optimizing team workflows
- Prefer automation over manual processes, fostering a culture of building software solutions to eliminate toil
- Deep technical familiarity with Kubernetes ecosystems, containerization technologies, and modern IaC tooling (e.g., Terraform, Crossplane, or Operators)
- Operated or supported stateful storage or database systems at scale and comfortable with durability, consistency and recovery trade-offs
- Excel at translating complex business and engineering requirements into actionable, phased technical roadmaps
- High level of empathy, responsibility, ownership, and accountability
- Excellent verbal and written technical communication skills
Nice-to-Haves
- Leading major architectural shifts, such as moving from legacy storage stacks to new multi-tenant storage architectures, including planning and executing large-scale data and workload migrations with tight availability and durability requirements
- Managing and scaling infrastructure across multi-cloud environments (AWS, GCP, or Azure)
- Designing secure, multi-tenant runtime environments at scale
Skills
Kubernetes, Terraform, Crossplane, AWS, GCP, Azure, Distributed Systems, Containerization, Iac, Storage Systems, MongoDB, Operators
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.
Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.
Build internal developer platforms, reusable services, and automation that improve software delivery, infrastructure self-service, reliability, and developer productivity. The role requires 5+ years of platform, software, infrastructure, DevOps, or SRE experience plus expertise in cloud-native technologies, Kubernetes, CI/CD, and Infrastructure as Code.