Staff Software Engineer, Observability
Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.
About the job
Responsibilities
- Create and maintain the foundational platform for running Reddit’s infrastructure.
- Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
- Contribute feedback to the technical and strategic direction of eventing at Reddit.
- Automate critical aspects of the event-driven development process.
- Share on-call responsibilities.
- Contribute upstream changes to the open-source projects used by the team.
Requirements
- 7+ years of experience developing internet-scale software, preferably in infrastructure.
- Familiarity with distributed systems development.
- Experience developing on Kubernetes or similar distributed systems.
- Strong troubleshooting capabilities across systems and software.
- Experience engineering large systems, tracking work, and independently driving projects.
- Excellent communication skills for collaboration with a service-oriented team and company.
Nice-to-haves
- Experience with Prometheus, Thanos, Grafana, Vector, ClickHouse, OpenTelemetry, or Loki.
- Kubernetes controller or operator development experience.
Compensation and Benefits
- Base salary range: $217,000–$303,900 USD.
- Equity in the form of restricted stock units may be available.
- Comprehensive healthcare benefits and income replacement programs.
- 401(k) with employer match.
- Global benefits supporting workspace, professional development, and caregiving.
- Family planning support.
- Gender-affirming care.
- Mental health and coaching benefits.
- Flexible vacation and paid volunteer time off.
- Paid parental leave.
Skills
Prometheus, Thanos, Grafana, Vector, ClickHouse, OpenTelemetry, Loki, Kubernetes, Distributed Systems, Kubernetes Operators
Similar jobs
DevOps / SRE jobsLeads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.
Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.