Skip to content
RedditReddit

Staff Site Reliability Engineer, Ads

Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.

About the job

Responsibilities

  • Lead reliability initiatives across Ads domains, including ad serving, auctions, targeting, reporting, measurement, and billing.
  • Partner with engineering leadership to develop a roadmap for reliability, scalability, operational excellence, and engineering efficiency.
  • Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
  • Drive architecture reviews and influence technical decisions affecting critical revenue-generating systems.
  • Participate in on-call rotations, lead complex incident investigations, and coordinate cross-functional responses during major production events.
  • Identify systemic reliability risks and drive long-term resilience improvements.
  • Establish reliability metrics for advertiser-critical journeys, including campaign creation, ad delivery, auction participation, reporting, attribution, and billing.
  • Mentor engineers and provide technical leadership across multiple teams.
  • Ensure reliability considerations are incorporated into product and infrastructure investments.

Requirements

  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
  • Experience evolving high-traffic, user-facing production environments.
  • Strong cross-functional collaboration and project leadership skills.
  • Deep expertise in distributed systems, scale engineering, and cloud-native architectures.
  • Experience designing highly available systems with strong operational and reliability practices.
  • Strong software engineering skills in general-purpose backend languages such as Go.
  • Understanding of observability systems, including metrics, logging, tracing, and alerting.
  • Experience improving reliability through SLOs, automation, incident management, and performance optimization.
  • Ability to troubleshoot complex issues across modern distributed system stacks.
  • Strong communication skills and ability to influence technical direction across teams.

Nice to Have

  • Experience supporting advertising technology platforms or other large-scale revenue-critical systems.
  • Understanding of reliability challenges in ad serving, real-time auctions, budget pacing, campaign delivery, measurement, attribution, or billing systems.
  • Experience operating high-QPS, low-latency services.
  • Experience establishing reliability programs with measurable business outcomes.
  • Experience with Kubernetes, cloud infrastructure, and large-scale distributed systems.
  • Familiarity with Kafka, ClickHouse, Spark, Flink, BigQuery, or similar large-scale data platforms.
  • Experience partnering with Product, Data Science, and Ads Engineering teams.
  • Experience supporting machine learning inference or recommendation systems at scale.

Compensation and Benefits

  • Base salary: $217,000–$303,900 USD.
  • Equity may be provided in the form of restricted stock units.
  • Health benefits, 401(k) matching, home-office benefits, professional development funds, family planning support, flexible vacation, global days off, paid parental leave, and paid volunteer time off.

Skills

Site Reliability Engineering, Distributed Systems, Go, Cloud-Native Architecture, Kubernetes, Observability, SLOs, Incident Management, Automation, Performance Optimization, Kafka, ClickHouse, Spark, Flink, BigQuery

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer - Site Experience
$217k+/yrOn-site8+ YOEDevOps / SRE

Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.

Reddit

Reddit

United States

Staff Software Engineer, Observability
$217k+/yrRemote7+ YOEDevOps / SRE

Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.

Coinbase

Coinbase

United States

Staff Software Engineer, Developer Infrastructure
$218k+/yrRemote8+ YOEDevOps / SRE

Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.

Coinbase

Coinbase

United States

Staff Infrastructure Engineer, Trading
$218k+/yrRemote8+ YOEDevOps / SRE

Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.

Crusoe

Crusoe

San Francisco, CA
Staff Software Engineer
$215k+/yrOn-site7+ YOEDevOps / SRE

Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.