Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.
217k – 304k/yr
Remote7+ YOEDevOps / SRE
About the role
Responsibilities
Create and maintain the foundational platform for running Reddit’s infrastructure.
Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
Contribute feedback to the technical and strategic direction of eventing at Reddit.
Automate critical aspects of the event-driven development process.
Share on-call responsibilities.
Contribute upstream changes to the open-source projects used by the team.
Requirements
7+ years of experience developing internet-scale software, preferably in infrastructure.
Familiarity with distributed systems development.
Experience developing on Kubernetes or similar distributed systems.
Strong troubleshooting capabilities across systems and software.
Experience engineering large systems, tracking work, and independently driving projects.
Excellent communication skills for collaboration with a service-oriented team and company.
Nice-to-haves
Experience with Prometheus, Thanos, Grafana, Vector, ClickHouse, OpenTelemetry, or Loki.
Kubernetes controller or operator development experience.
Compensation and Benefits
Base salary range: $217,000–$303,900 USD.
Equity in the form of restricted stock units may be available.
Comprehensive healthcare benefits and income replacement programs.
401(k) with employer match.
Global benefits supporting workspace, professional development, and caregiving.
Staff Software Engineer on the Developer Experience team at Grow Therapy, owning high-impact platform and AI-native tooling to accelerate 120+ engineers. Sets org-wide standards for build systems, CI/CD, monorepo performance, and agentic AI workflows while driving monolith decomposition and measuring adoption.
217k – 289k/yrHybrid7+ YOEDevOps / SRE
Staff Site Reliability Engineer
IdmeMountain View, CA
Leads infrastructure transformation from monoliths to scalable microservices at massive scale, architects observability/CI/CD systems, unifies complex stacks, and mentors engineers. Requires 10+ years coding internal tools, 5+ years cloud (GCP/AWS), Bachelor's in CS.
218k – 260k/yrOn-site10+ YOEDevOps / SRE
Staff Infrastructure Software Engineer, Enterprise AI
Scale AINew York, NY +1
Builds and scales multi-cloud infrastructure for enterprise AI Agentic workflows, focusing on security, compliance, observability, and developer tools. Requires 5+ years experience with modern infra practices, cloud providers, and languages like Python.
216k – 270k/yrHybrid5+ YOEDevOps / SRE
Staff Software Engineer
CoinbaseUnited States
Staff Software Engineer owning technical strategy and systems for Coinbase's test infrastructure at scale. Focus on fast, reliable test signals through orchestration, smart selection, sharding, and flakiness remediation.
218k – 257k/yrRemote10+ YOEDevOps / SRE
Staff Site Reliability Engineer
CoinbaseUnited States
Staff SRE on the IT Operations team owning reliability, automation, and observability for Coinbase's AI infrastructure on AWS and Kubernetes. Requires 8+ years of cloud infrastructure experience and strong incident response leadership.