Skip to content
RedditRedditUnited States

Staff Software Engineer, Observability

Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.

217k – 304k/yr
Remote7+ YOEDevOps / SRE

About the role

Responsibilities

  • Create and maintain the foundational platform for running Reddit’s infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute feedback to the technical and strategic direction of eventing at Reddit.
  • Automate critical aspects of the event-driven development process.
  • Share on-call responsibilities.
  • Contribute upstream changes to the open-source projects used by the team.

Requirements

  • 7+ years of experience developing internet-scale software, preferably in infrastructure.
  • Familiarity with distributed systems development.
  • Experience developing on Kubernetes or similar distributed systems.
  • Strong troubleshooting capabilities across systems and software.
  • Experience engineering large systems, tracking work, and independently driving projects.
  • Excellent communication skills for collaboration with a service-oriented team and company.

Nice-to-haves

  • Experience with Prometheus, Thanos, Grafana, Vector, ClickHouse, OpenTelemetry, or Loki.
  • Kubernetes controller or operator development experience.

Compensation and Benefits

  • Base salary range: $217,000–$303,900 USD.
  • Equity in the form of restricted stock units may be available.
  • Comprehensive healthcare benefits and income replacement programs.
  • 401(k) with employer match.
  • Global benefits supporting workspace, professional development, and caregiving.
  • Family planning support.
  • Gender-affirming care.
  • Mental health and coaching benefits.
  • Flexible vacation and paid volunteer time off.
  • Paid parental leave.

Skills

PrometheusthanosGrafanavectorClickHouseOpenTelemetrylokiKubernetesDistributed Systemskubernetes operators

Similar roles

DevOps / SRE jobs
Grow Therapy

Staff Software Engineer

Grow TherapySan Francisco, CA +1

Staff Software Engineer on the Developer Experience team at Grow Therapy, owning high-impact platform and AI-native tooling to accelerate 120+ engineers. Sets org-wide standards for build systems, CI/CD, monorepo performance, and agentic AI workflows while driving monolith decomposition and measuring adoption.

217k – 289k/yrHybrid7+ YOEDevOps / SRE
Idme

Staff Site Reliability Engineer

IdmeMountain View, CA

Leads infrastructure transformation from monoliths to scalable microservices at massive scale, architects observability/CI/CD systems, unifies complex stacks, and mentors engineers. Requires 10+ years coding internal tools, 5+ years cloud (GCP/AWS), Bachelor's in CS.

218k – 260k/yrOn-site10+ YOEDevOps / SRE
Scale AI

Staff Infrastructure Software Engineer, Enterprise AI

Scale AINew York, NY +1

Builds and scales multi-cloud infrastructure for enterprise AI Agentic workflows, focusing on security, compliance, observability, and developer tools. Requires 5+ years experience with modern infra practices, cloud providers, and languages like Python.

216k – 270k/yrHybrid5+ YOEDevOps / SRE
Coinbase

Staff Software Engineer

CoinbaseUnited States

Staff Software Engineer owning technical strategy and systems for Coinbase's test infrastructure at scale. Focus on fast, reliable test signals through orchestration, smart selection, sharding, and flakiness remediation.

218k – 257k/yrRemote10+ YOEDevOps / SRE
Coinbase

Staff Site Reliability Engineer

CoinbaseUnited States

Staff SRE on the IT Operations team owning reliability, automation, and observability for Coinbase's AI infrastructure on AWS and Kubernetes. Requires 8+ years of cloud infrastructure experience and strong incident response leadership.

218k – 257k/yrRemote8+ YOEDevOps / SRE