Staff Site Reliability Engineer - Site Experience
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
About the job
Responsibilities
- Lead reliability engineering for critical user-facing systems and services across APIs, content delivery, feed generation, search, messaging, and real-time experiences.
- Improve availability, latency, scalability, performance, and resiliency under large-scale global load.
- Guide architecture decisions involving failover, redundancy, graceful degradation, traffic management, and capacity planning.
- Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure.
- Build automation and tooling for deployment safety, incident response, remediation workflows, and reliability guardrails.
- Lead complex incident response, blameless postmortems, root-cause analysis, and sustainable remediation.
- Define and promote practices for reliability engineering, SLIs/SLOs, capacity management, release engineering, and operational maturity.
- Provide technical leadership and mentorship across SRE and software engineering teams.
Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
- Experience supporting high-traffic, user-facing production environments.
- Deep knowledge of distributed systems, networking, Linux systems, or cloud-native architectures.
- Experience designing highly available systems and applying strong operational and reliability practices.
- Strong programming skills in Go, Python, or similar languages.
- Understanding of observability systems, including metrics, logging, tracing, and alerting.
- Experience improving reliability through SLOs, automation, incident management, and performance optimization.
- Ability to troubleshoot complex issues across applications, infrastructure, networking, and services.
- Strong collaboration and communication skills, with the ability to influence technical direction across teams.
Nice to Have
- Experience operating systems at internet-scale traffic volumes.
- Experience with Kubernetes, containers, cloud infrastructure, and modern deployment platforms.
- Familiarity with Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis, or similar distributed infrastructure technologies.
- Experience with CDN optimization, edge reliability, traffic engineering, or global infrastructure.
- Contributions to open source software or participation in technical communities.
- Experience leading large-scale incident response and operational transformation initiatives.
Compensation
- Base salary: $217,000–$303,900 USD.
- Eligible for equity in the form of restricted stock units; select positions may also be eligible for commission.
- Benefits may include medical, dental, and vision insurance, 401(k) matching, paid vacation, volunteer time off, parental leave, family planning support, gender-affirming care, mental health and coaching benefits, professional development support, and caregiving support.
Skills
Go, Python, Distributed Systems, Networking, Linux, Cloud-Native Architecture, Kubernetes, Containers, Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra
Similar jobs
DevOps / SRE jobsProvides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.
Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.
Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.