# Staff Site Reliability Engineer, Ads

**Company:** [Reddit](https://hotfix.jobs/companies/reddit)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $217k – $304k/yr
**Experience:** 8+ years
**Skills:** Site Reliability Engineering, Distributed Systems, Go, Cloud-Native Architecture, Kubernetes, Observability, SLOs, Incident Management, Automation, Performance Optimization, Kafka, ClickHouse, Spark, Flink, BigQuery
**Posted:** 2026-08-29

> Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.

## Job Description

## Responsibilities
- Lead reliability initiatives across Ads domains, including ad serving, auctions, targeting, reporting, measurement, and billing.
- Partner with engineering leadership to develop a roadmap for reliability, scalability, operational excellence, and engineering efficiency.
- Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
- Drive architecture reviews and influence technical decisions affecting critical revenue-generating systems.
- Participate in on-call rotations, lead complex incident investigations, and coordinate cross-functional responses during major production events.
- Identify systemic reliability risks and drive long-term resilience improvements.
- Establish reliability metrics for advertiser-critical journeys, including campaign creation, ad delivery, auction participation, reporting, attribution, and billing.
- Mentor engineers and provide technical leadership across multiple teams.
- Ensure reliability considerations are incorporated into product and infrastructure investments.

## Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
- Experience evolving high-traffic, user-facing production environments.
- Strong cross-functional collaboration and project leadership skills.
- Deep expertise in distributed systems, scale engineering, and cloud-native architectures.
- Experience designing highly available systems with strong operational and reliability practices.
- Strong software engineering skills in general-purpose backend languages such as Go.
- Understanding of observability systems, including metrics, logging, tracing, and alerting.
- Experience improving reliability through SLOs, automation, incident management, and performance optimization.
- Ability to troubleshoot complex issues across modern distributed system stacks.
- Strong communication skills and ability to influence technical direction across teams.

## Nice to Have
- Experience supporting advertising technology platforms or other large-scale revenue-critical systems.
- Understanding of reliability challenges in ad serving, real-time auctions, budget pacing, campaign delivery, measurement, attribution, or billing systems.
- Experience operating high-QPS, low-latency services.
- Experience establishing reliability programs with measurable business outcomes.
- Experience with Kubernetes, cloud infrastructure, and large-scale distributed systems.
- Familiarity with Kafka, ClickHouse, Spark, Flink, BigQuery, or similar large-scale data platforms.
- Experience partnering with Product, Data Science, and Ads Engineering teams.
- Experience supporting machine learning inference or recommendation systems at scale.

## Compensation and Benefits
- Base salary: **$217,000–$303,900 USD**.
- Equity may be provided in the form of restricted stock units.
- Health benefits, 401(k) matching, home-office benefits, professional development funds, family planning support, flexible vacation, global days off, paid parental leave, and paid volunteer time off.

## Similar jobs

- [Staff Site Reliability Engineer - Site Experience](https://hotfix.jobs/jobs/d8fa80d2-f6f6-4a20-8485-60d2648f8daf) - Reddit - San Francisco, CA - $217k – $304k/yr
- [Staff Software Engineer, Observability](https://hotfix.jobs/jobs/002663cc-6743-4d3d-9bd8-a1e5c012ed86) - Reddit - Remote - $217k – $304k/yr
- [Staff Software Engineer, Developer Infrastructure](https://hotfix.jobs/jobs/9e5df28a-b02e-49c1-ae34-8ac4874bd491) - Coinbase - Remote - $218k – $257k/yr
- [Staff Infrastructure Engineer, Trading](https://hotfix.jobs/jobs/4be49748-b215-44fa-b91e-9975c38847a9) - Coinbase - Remote - $218k – $257k/yr
- [Staff Software Engineer](https://hotfix.jobs/jobs/ee5e02a2-cdbb-433e-88e7-c91608e76e93) - Crusoe - San Francisco, CA - $215k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/91c6ec8f-96a4-469b-ad89-c0a0aeb77e26
**Canonical:** https://hotfix.jobs/jobs/91c6ec8f-96a4-469b-ad89-c0a0aeb77e26