# Staff Site Reliability Engineer - Site Experience

**Company:** [Reddit](https://hotfix.jobs/companies/reddit)
**Location:** San Francisco, CA
**Role:** DevOps / SRE
**Salary:** $217k – $304k/yr
**Experience:** 8+ years
**Skills:** Go, Python, Distributed Systems, Networking, Linux, Cloud-Native Architecture, Kubernetes, Containers, Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra
**Posted:** 2026-09-01

> Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.

## Job Description

## Responsibilities
- Lead reliability engineering for critical user-facing systems and services across APIs, content delivery, feed generation, search, messaging, and real-time experiences.
- Improve availability, latency, scalability, performance, and resiliency under large-scale global load.
- Guide architecture decisions involving failover, redundancy, graceful degradation, traffic management, and capacity planning.
- Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure.
- Build automation and tooling for deployment safety, incident response, remediation workflows, and reliability guardrails.
- Lead complex incident response, blameless postmortems, root-cause analysis, and sustainable remediation.
- Define and promote practices for reliability engineering, SLIs/SLOs, capacity management, release engineering, and operational maturity.
- Provide technical leadership and mentorship across SRE and software engineering teams.

## Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
- Experience supporting high-traffic, user-facing production environments.
- Deep knowledge of distributed systems, networking, Linux systems, or cloud-native architectures.
- Experience designing highly available systems and applying strong operational and reliability practices.
- Strong programming skills in Go, Python, or similar languages.
- Understanding of observability systems, including metrics, logging, tracing, and alerting.
- Experience improving reliability through SLOs, automation, incident management, and performance optimization.
- Ability to troubleshoot complex issues across applications, infrastructure, networking, and services.
- Strong collaboration and communication skills, with the ability to influence technical direction across teams.

## Nice to Have
- Experience operating systems at internet-scale traffic volumes.
- Experience with Kubernetes, containers, cloud infrastructure, and modern deployment platforms.
- Familiarity with Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis, or similar distributed infrastructure technologies.
- Experience with CDN optimization, edge reliability, traffic engineering, or global infrastructure.
- Contributions to open source software or participation in technical communities.
- Experience leading large-scale incident response and operational transformation initiatives.

## Compensation
- Base salary: $217,000–$303,900 USD.
- Eligible for equity in the form of restricted stock units; select positions may also be eligible for commission.
- Benefits may include medical, dental, and vision insurance, 401(k) matching, paid vacation, volunteer time off, parental leave, family planning support, gender-affirming care, mental health and coaching benefits, professional development support, and caregiving support.

## Similar jobs

- [Staff Site Reliability Engineer, Ads](https://hotfix.jobs/jobs/91c6ec8f-96a4-469b-ad89-c0a0aeb77e26) - Reddit - Remote - $217k – $304k/yr
- [Staff Software Engineer, Observability](https://hotfix.jobs/jobs/002663cc-6743-4d3d-9bd8-a1e5c012ed86) - Reddit - Remote - $217k – $304k/yr
- [Staff Software Engineer, Developer Infrastructure](https://hotfix.jobs/jobs/9e5df28a-b02e-49c1-ae34-8ac4874bd491) - Coinbase - Remote - $218k – $257k/yr
- [Staff Infrastructure Engineer, Trading](https://hotfix.jobs/jobs/4be49748-b215-44fa-b91e-9975c38847a9) - Coinbase - Remote - $218k – $257k/yr
- [Staff Software Engineer](https://hotfix.jobs/jobs/ee5e02a2-cdbb-433e-88e7-c91608e76e93) - Crusoe - San Francisco, CA - $215k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/d8fa80d2-f6f6-4a20-8485-60d2648f8daf
**Canonical:** https://hotfix.jobs/jobs/d8fa80d2-f6f6-4a20-8485-60d2648f8daf