Skip to content
MongoDBMongoDB

Site Reliability Engineer , Storage Layer Services

This SRE will operate and improve MongoDB Atlas’s multi-tenant distributed storage infrastructure, focusing on reliability, performance, observability, automation, and incident response. The role requires 6+ years of distributed-systems experience plus expertise in storage or databases, Kubernetes, cloud platforms, Linux, and networking.

About the job

Responsibilities

  • Work on multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs.
  • Build reliable services and infrastructure that are available, resilient, fault-tolerant, and self-healing.
  • Identify and configure key metrics to detect incidents and quantify service health, availability, and performance.
  • Participate in a 24/7 on-call rotation to resolve issues involving storage infrastructure.
  • Become an expert in infrastructure performance, optimizing from the application level through the kernel.

Requirements

  • 6+ years of experience working on software development and operating distributed systems.
  • Proficiency in Python, Go, or a similar language.
  • Experience operating or supporting stateful storage or database systems at scale, including durability, consistency, and recovery trade-offs.
  • Customer-focused mindset and an emphasis on operational efficiency.
  • Preference for automation over manual processes and reducing operational toil through software solutions.
  • Experience using and extending containerization technologies, particularly Kubernetes.
  • Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform, or Azure.
  • Understanding of Linux operating system internals and networking concepts such as TCP/IP, DNS, TLS, and routing.

Nice-to-haves

  • Experience leading major architectural shifts, including migrations from legacy storage stacks to new multi-tenant storage architectures.
  • Experience planning and executing large-scale data and workload migrations with strict availability and durability requirements.
  • Experience managing and scaling infrastructure across AWS, Google Cloud Platform, or Azure.
  • Experience designing secure, multi-tenant runtime environments at scale.

Skills

Python, Go, Kubernetes, AWS, GCP, Microsoft Azure, Linux, TCP/IP, DNS, Tls, Routing, Distributed Systems, Database Systems, Containerization

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Muck Rack

Muck Rack

Bulgaria
Senior Software Engineer, DevOps
€95k+/yrRemote5+ YOEDevOps / SRE

Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.

Kraken

Kraken

United Arab Emirates
Senior Database Administrator - Core Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Crusoe

Crusoe

Dublin, Ireland

Senior Network Production Operations Engineer
No salary listedOn-site10+ YOEDevOps / SRE

Operates and scales Crusoe Cloud’s global edge, backbone, and data center networks supporting GPU-based HPC workloads. The role requires extensive production networking experience, strong protocol and observability expertise, automation skills, and participation in 24/7 on-call support.