Skip to content
RedditReddit

Staff Software Engineer, Observability

Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.

About the job

Responsibilities

  • Create and maintain the foundational platform for running Reddit’s infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute feedback to the technical and strategic direction of eventing at Reddit.
  • Automate critical aspects of the event-driven development process.
  • Share on-call responsibilities.
  • Contribute upstream changes to the open-source projects used by the team.

Requirements

  • 7+ years of experience developing internet-scale software, preferably in infrastructure.
  • Familiarity with distributed systems development.
  • Experience developing on Kubernetes or similar distributed systems.
  • Strong troubleshooting capabilities across systems and software.
  • Experience engineering large systems, tracking work, and independently driving projects.
  • Excellent communication skills for collaboration with a service-oriented team and company.

Nice-to-haves

  • Experience with Prometheus, Thanos, Grafana, Vector, ClickHouse, OpenTelemetry, or Loki.
  • Kubernetes controller or operator development experience.

Compensation and Benefits

  • Base salary range: $217,000–$303,900 USD.
  • Equity in the form of restricted stock units may be available.
  • Comprehensive healthcare benefits and income replacement programs.
  • 401(k) with employer match.
  • Global benefits supporting workspace, professional development, and caregiving.
  • Family planning support.
  • Gender-affirming care.
  • Mental health and coaching benefits.
  • Flexible vacation and paid volunteer time off.
  • Paid parental leave.

Skills

Prometheus, Thanos, Grafana, Vector, ClickHouse, OpenTelemetry, Loki, Kubernetes, Distributed Systems, Kubernetes Operators

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer - Site Experience
$217k+/yrOn-site8+ YOEDevOps / SRE

Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer, Ads
$217k+/yrRemote8+ YOEDevOps / SRE

Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.

Coinbase

Coinbase

United States

Staff Software Engineer, Developer Infrastructure
$218k+/yrRemote8+ YOEDevOps / SRE

Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.

Coinbase

Coinbase

United States

Staff Infrastructure Engineer, Trading
$218k+/yrRemote8+ YOEDevOps / SRE

Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.

Crusoe

Crusoe

San Francisco, CA
Staff Software Engineer
$215k+/yrOn-site7+ YOEDevOps / SRE

Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.