Latest DevOps / SRE jobs at MongoDB
Job results
Senior network engineer responsible for designing, operating, and securing MongoDB’s global network and VPN infrastructure. The role requires 6+ years of networking or systems engineering experience, strong enterprise networking expertise, automation skills, and the ability to lead complex infrastructure initiatives.
Provides technical leadership for the reliability architecture and operational foundations of a multi-region, multi-cloud platform for AI applications. The role requires deep Kubernetes and distributed-systems expertise, infrastructure programming, cloud knowledge, and experience setting SRE standards and mentoring engineers.
Build and operate a self-service internal development platform that helps engineering teams deploy and run production services reliably. The role requires strong backend programming, production Kubernetes operations, cloud infrastructure, observability, networking, and distributed-systems experience.
The Senior Site Reliability Engineer will operate and improve Kubernetes-based distributed infrastructure for AI application workloads, focusing on scalability, observability, reliability, and tenant isolation. The role requires 6+ years of distributed-systems experience, production Kubernetes expertise, cloud infrastructure knowledge, and strong programming skills.
Leads Cloud Operations Engineering activities, combining technical leadership, incident response, production troubleshooting, automation, and team coaching. The role requires expertise in Linux, networking, cloud infrastructure, monitoring, distributed systems, and at least two programming languages.
Staff Engineer responsible for architecting and operating MongoDB’s large-scale observability collection and ingestion infrastructure. The role requires 10+ years of experience with distributed or highly concurrent systems, expert programming skills, and strong database, performance, and systems fundamentals.
This SRE will operate and improve MongoDB Atlas’s multi-tenant distributed storage infrastructure, focusing on reliability, performance, observability, automation, and incident response. The role requires 6+ years of distributed-systems experience plus expertise in storage or databases, Kubernetes, cloud platforms, Linux, and networking.
Cloud Operations Engineer on 2nd shift weekends responsible for monitoring Atlas platform, diagnosing incidents, on-call rotations, automation, and ensuring uptime for MongoDB customers in FedRamp environments. Requires 2+ years DevOps/SRE experience, Linux expertise, cloud familiarity, and scripting skills.
Senior or Staff Site Reliability Engineer focused on continuous delivery infrastructure using Argo Workflows, ArgoCD, and Kubernetes. Owns deployment tooling, onboarding flows, and participates in 24/7 on-call. Requires 6+ years building and operating distributed systems.
Designs, builds, and optimizes global infrastructure for MongoDB Atlas, focusing on automation, monitoring, resilience, and performance at massive scale. Requires 3+ years experience with Linux services, programming, and automation tools.
Senior or Staff Site Reliability Engineer maintains and scales the Atlas platform in a multi-cloud environment, focusing on automation, on-call reliability, and collaboration with engineering teams. Requires 5+ years experience with Linux, cloud providers, and programming languages like Go, Python, or Ruby.
Senior SRE ensures reliability of MongoDB's multi-tenant cloud storage layer by defining SLOs, building resilient infrastructure, optimizing performance, and participating in on-call. Requires 6+ years experience with distributed systems, Python/Go, Kubernetes, and cloud platforms.
Staff SRE on the Fabric team builds and maintains secure multi-cloud networking infrastructure for service communication, leveraging deep networking expertise to ensure resilience and scalability. Requires 10+ years experience in distributed systems and networking fundamentals.
Senior or Staff Site Reliability Engineer leads design and implementation of cloud security solutions (AWS, Azure, GCP), builds automation for monitoring and alerting, and mentors SRE team. Requires 6+ years SRE/infra experience with security focus, IaC proficiency, and cloud expertise.
Leads a team of SREs for MongoDB's Storage Layer Services, defining SLOs, capacity plans, and roadmaps for multi-tenant distributed storage systems underpinning Atlas. Requires 10+ years in distributed systems and 2+ years managing teams, with expertise in Kubernetes and IaC tools.
Senior SRE on the Fleet Management team develops and maintains scalable Kubernetes runtime environments, provides internal support to engineering teams, and participates in 24/7 on-call with a focus on automation and blameless post-mortems. Requires 6+ years experience with distributed systems, containerization, and Go/Python proficiency.