Skip to content
MongoDBMongoDB

Senior Site Reliability Engineer

Senior SRE ensures reliability of MongoDB's multi-tenant cloud storage layer by defining SLOs, building resilient infrastructure, optimizing performance, and participating in on-call. Requires 6+ years experience with distributed systems, Python/Go, Kubernetes, and cloud platforms.

About the job

Responsibilities

  • Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs
  • Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing
  • Identify and configure key metrics to detect incidents and quantify service health, availability, and performance
  • Participate in a 24/7 on-call rotation to resolve issues involving the storage infrastructure
  • Become an expert in infrastructure performance, helping us optimize from the application level all the way to the kernel

Requirements

  • 6+ years of experience working on software development and operating distributed systems
  • Proficiency in Python, Go, or a similar language
  • Operated or supported stateful storage or database systems at scale, comfortable with durability, consistency, and recovery trade-offs
  • Customer-focused mindset
  • Value efficiency in processes and operations
  • Prefer automation over manual processes
  • Experience using and extending containerization technologies, particularly Kubernetes
  • Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure
  • Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing)

Nice-to-Haves

  • Leading major architectural shifts, such as moving from legacy storage stacks to new multi-tenant storage architectures, including planning and executing large-scale data and workload migrations with tight availability and durability requirements
  • Managing and scaling infrastructure across multi-cloud environments (AWS, GCP, or Azure)
  • Designing secure, multi-tenant runtime environments at scale

Skills

Python, Go, Kubernetes, AWS, GCP, Azure, Linux, Distributed Systems, Storage Systems, MongoDB

PrizePicks

PrizePicks

United States

Senior Site Reliability Engineer
$120k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for designing and operating reliable, scalable production infrastructure, leading incident response, and improving observability and resilience. Requires 5+ years of reliability-focused engineering experience and expertise across cloud, infrastructure as code, Kubernetes, monitoring, and application development.

Axle

Axle

Rockville, MD

Lead Scientific Imaging Systems Engineer
$120k+/yrOn-site7+ YOEDevOps / SRE

Leads end-to-end infrastructure for a scientific imaging platform, covering Linux administration, GPU/HPC systems, storage, upgrades, and vendor coordination. The role supports AI-enabled imaging workflows and requires extensive production Linux, Image Artist, GPU, HPC, and enterprise storage experience.

MongoDB

MongoDB

Palo Alto, CA
Senior Network Engineer
$118k+/yrHybrid6+ YOEDevOps / SRE

Senior network engineer responsible for designing, operating, and securing MongoDB’s global network and VPN infrastructure. The role requires 6+ years of networking or systems engineering experience, strong enterprise networking expertise, automation skills, and the ability to lead complex infrastructure initiatives.

Upstart

Upstart

United States

Senior DevOps Engineer
$136k+/yrRemote3+ YOEDevOps / SRE

Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.

Bloomerang

Bloomerang

United States

Senior Software Engineer, Site Reliability
$115k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.