Skip to content
RedditReddit

Staff Site Reliability Engineer - Site Experience

Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.

About the job

Responsibilities

  • Lead reliability engineering for critical user-facing systems and services across APIs, content delivery, feed generation, search, messaging, and real-time experiences.
  • Improve availability, latency, scalability, performance, and resiliency under large-scale global load.
  • Guide architecture decisions involving failover, redundancy, graceful degradation, traffic management, and capacity planning.
  • Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure.
  • Build automation and tooling for deployment safety, incident response, remediation workflows, and reliability guardrails.
  • Lead complex incident response, blameless postmortems, root-cause analysis, and sustainable remediation.
  • Define and promote practices for reliability engineering, SLIs/SLOs, capacity management, release engineering, and operational maturity.
  • Provide technical leadership and mentorship across SRE and software engineering teams.

Requirements

  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
  • Experience supporting high-traffic, user-facing production environments.
  • Deep knowledge of distributed systems, networking, Linux systems, or cloud-native architectures.
  • Experience designing highly available systems and applying strong operational and reliability practices.
  • Strong programming skills in Go, Python, or similar languages.
  • Understanding of observability systems, including metrics, logging, tracing, and alerting.
  • Experience improving reliability through SLOs, automation, incident management, and performance optimization.
  • Ability to troubleshoot complex issues across applications, infrastructure, networking, and services.
  • Strong collaboration and communication skills, with the ability to influence technical direction across teams.

Nice to Have

  • Experience operating systems at internet-scale traffic volumes.
  • Experience with Kubernetes, containers, cloud infrastructure, and modern deployment platforms.
  • Familiarity with Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis, or similar distributed infrastructure technologies.
  • Experience with CDN optimization, edge reliability, traffic engineering, or global infrastructure.
  • Contributions to open source software or participation in technical communities.
  • Experience leading large-scale incident response and operational transformation initiatives.

Compensation

  • Base salary: $217,000–$303,900 USD.
  • Eligible for equity in the form of restricted stock units; select positions may also be eligible for commission.
  • Benefits may include medical, dental, and vision insurance, 401(k) matching, paid vacation, volunteer time off, parental leave, family planning support, gender-affirming care, mental health and coaching benefits, professional development support, and caregiving support.

Skills

Go, Python, Distributed Systems, Networking, Linux, Cloud-Native Architecture, Kubernetes, Containers, Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer, Ads
$217k+/yrRemote8+ YOEDevOps / SRE

Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.

Reddit

Reddit

United States

Staff Software Engineer, Observability
$217k+/yrRemote7+ YOEDevOps / SRE

Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.

Coinbase

Coinbase

United States

Staff Software Engineer, Developer Infrastructure
$218k+/yrRemote8+ YOEDevOps / SRE

Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.

Coinbase

Coinbase

United States

Staff Infrastructure Engineer, Trading
$218k+/yrRemote8+ YOEDevOps / SRE

Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.

Crusoe

Crusoe

San Francisco, CA
Staff Software Engineer
$215k+/yrOn-site7+ YOEDevOps / SRE

Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.