Staff Site Reliability Engineer, Ads
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.
About the job
Responsibilities
- Lead reliability initiatives across Ads domains, including ad serving, auctions, targeting, reporting, measurement, and billing.
- Partner with engineering leadership to develop a roadmap for reliability, scalability, operational excellence, and engineering efficiency.
- Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
- Drive architecture reviews and influence technical decisions affecting critical revenue-generating systems.
- Participate in on-call rotations, lead complex incident investigations, and coordinate cross-functional responses during major production events.
- Identify systemic reliability risks and drive long-term resilience improvements.
- Establish reliability metrics for advertiser-critical journeys, including campaign creation, ad delivery, auction participation, reporting, attribution, and billing.
- Mentor engineers and provide technical leadership across multiple teams.
- Ensure reliability considerations are incorporated into product and infrastructure investments.
Requirements
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
- Experience evolving high-traffic, user-facing production environments.
- Strong cross-functional collaboration and project leadership skills.
- Deep expertise in distributed systems, scale engineering, and cloud-native architectures.
- Experience designing highly available systems with strong operational and reliability practices.
- Strong software engineering skills in general-purpose backend languages such as Go.
- Understanding of observability systems, including metrics, logging, tracing, and alerting.
- Experience improving reliability through SLOs, automation, incident management, and performance optimization.
- Ability to troubleshoot complex issues across modern distributed system stacks.
- Strong communication skills and ability to influence technical direction across teams.
Nice to Have
- Experience supporting advertising technology platforms or other large-scale revenue-critical systems.
- Understanding of reliability challenges in ad serving, real-time auctions, budget pacing, campaign delivery, measurement, attribution, or billing systems.
- Experience operating high-QPS, low-latency services.
- Experience establishing reliability programs with measurable business outcomes.
- Experience with Kubernetes, cloud infrastructure, and large-scale distributed systems.
- Familiarity with Kafka, ClickHouse, Spark, Flink, BigQuery, or similar large-scale data platforms.
- Experience partnering with Product, Data Science, and Ads Engineering teams.
- Experience supporting machine learning inference or recommendation systems at scale.
Compensation and Benefits
- Base salary: $217,000–$303,900 USD.
- Equity may be provided in the form of restricted stock units.
- Health benefits, 401(k) matching, home-office benefits, professional development funds, family planning support, flexible vacation, global days off, paid parental leave, and paid volunteer time off.
Skills
Site Reliability Engineering, Distributed Systems, Go, Cloud-Native Architecture, Kubernetes, Observability, SLOs, Incident Management, Automation, Performance Optimization, Kafka, ClickHouse, Spark, Flink, BigQuery
Similar jobs
DevOps / SRE jobsLeads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.
Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.