Skip to content
MongoDBMongoDB

Staff Site Reliability Engineer, Fabric

Staff SRE on the Fabric team builds and maintains secure multi-cloud networking infrastructure for service communication, leveraging deep networking expertise to ensure resilience and scalability. Requires 10+ years experience in distributed systems and networking fundamentals.

About the job

Responsibilities

  • Participate in the development of a reliable and resilient multi-cloud globally-connected network that is crucial for MongoDB’s services.
  • Collaborate with service-owning teams to provide internal support, addressing technical issues and offering guidance on best practices for service-to-service connectivity.
  • Participate in a 24/7 on-call rotation to swiftly resolve issues related to network architecture and service-to-service connectivity, ensuring minimal disruption and high availability.

Requirements

  • 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles.
  • Customer-focused mindset, driving improvements that benefit end-users.
  • Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”).
  • Intimately familiar with modern cloud-based infrastructure and the network design primitives of at least one of AWS, Azure, or GCP, e.g. VPCs, subnetting, routing, VPNs, peering, private link / private service connect, and CDNs.
  • Strong knowledge of service mesh and load-balancing concepts, and be eager to implement these in a multi-cloud environment.

Compensation

Base salary range: $127,000—$249,000 USD. Other benefits may include equity, employee stock purchase program, flexible PTO, parental leave, 401(k), mental health counseling, and health insurance.

Skills

Kubernetes, TCP/IP, DNS, Tls, BGP, AWS, Azure, GCP, Service Mesh, Load Balancing, Ipv6, Sdn, Vpcs, Cdns

GitLab

GitLab

Canada
Site Reliability Engineer, Intermediate to Senior Staff
$126k+/yrRemote5+ YOEDevOps / SRE

Site Reliability Engineers build and operate scalable production infrastructure, automate operational workflows, and improve observability, incident response, and service reliability. The role spans Intermediate through Senior Staff levels and requires experience with Kubernetes, infrastructure as code, cloud platforms, and software engineering.

Nango

Nango

United States
Staff Engineer, Platform & Infrastructure
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.

Nango

Nango

United States
Staff Platform Engineer
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.

VGS

VGS

United States
Staff Infrastructure Engineer
$145k+/yrRemote8+ YOEDevOps / SRE

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting mission-critical payment systems. Requires 8+ years of distributed-systems experience and deep expertise in infrastructure as code, Kubernetes, automation, and cloud networking.

Mozilla

Mozilla

Canada

Senior Staff Performance Engineer, Firefox
CA$149k+/yrRemote7+ YOEDevOps / SRE

Leads Firefox performance engineering by writing code, profiling bottlenecks, improving benchmarks, and guiding cross-functional teams. Requires 7+ years of experience, strong C++ and JavaScript skills, and expertise in performance-critical software, profiling, concurrency, and systems analysis.