Skip to content
CrusoeCrusoe

Staff Software Engineer, Cloud Monitoring Service

Lead the design and evolution of Crusoe Cloud's large-scale telemetry and observability systems for metrics and logs. Own high-throughput distributed pipelines from edge collection through ingestion, storage, and low-latency querying while ensuring scalability, reliability, and multi-tenancy.

About the job

What You'll Be Working On

Distributed Systems Ownership: Own the architecture and evolution of large-scale telemetry pipelines, including high-volume ingestion, stream processing, time-series and log storage, and low-latency query paths. Design systems that stay correct, available, and cost-efficient as data volume grows 10x.

Scalable, Multi-Tenant Design: Design services that are highly scalable, durable, and fair across tenants. Solve the hard problems in this space: hot shards, high-cardinality data, noisy neighbors, backpressure, retention and compaction at scale, and graceful degradation under load.

Reliability and Operational Excellence: Build for operability from day one. Improve pipeline reliability and data freshness, reduce on-call burden through better system design rather than more process, and participate in a customer-facing on-call rotation, leading by example.

Technical Leadership: Set the technical direction for the team's distributed systems work. Drive design reviews, identify one-way door decisions early, and raise the bar on how the team scopes, builds, and operates systems.

Cross-Team Collaboration: Work with product, compute, networking, and platform teams to make sure observability decisions are made with full context. Represent the team's technical position in cross-org conversations.

Mentorship: Coach senior and mid-level engineers through design work, code review, and incident response. Build patterns and frameworks that make the team better without requiring your direct involvement.

What You'll Bring to the Team

Distributed Systems Depth: Deep, hands-on experience designing and operating distributed systems at scale. You have solved real problems in sharding, replication, consistency, load balancing, and concurrency, not just studied them.

Observability Data Experience: Experience building or operating large-scale data infrastructure such as time-series databases, log aggregation, streaming pipelines, or distributed tracing backends. Familiarity with technologies like Prometheus, VictoriaMetrics, Loki, OpenTelemetry, Kafka, Vector, or similar.

Technical Proficiency: Strong programming fundamentals in Go or another modern compiled language (Go strongly preferred). Comfort with Kubernetes, microservices, and CI/CD as the operating environment for everything you build.

Design Judgment at Staff Level: You proactively scope ambiguous problems, surface non-functional requirements without being prompted, and reason about tradeoffs in terms of customer and business outcomes, not just technical elegance.

Operational Mindset: On-call experience on a customer-facing service. You treat incidents as design feedback and systematically eliminate the conditions that cause them.

Cross-Functional Influence: A track record of driving technical outcomes that span team boundaries, and the communication skills to bring product, support, and adjacent engineering teams along.

Mentorship: You make the engineers around you better through design guidance, thoughtful review, and honest, constructive feedback.

Professional Experience: 8+ years of software development experience, with sustained ownership of production distributed systems.

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off, paid holidays & leave of absence programs
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance & emergency assistance
  • Daily meals allowance
  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $215,000 - $260,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Skills

Distributed Systems, Go, Kubernetes, Prometheus, Victoriametrics, Loki, OpenTelemetry, Kafka, Vector, Time-Series Databases, Log Aggregation, Streaming Pipelines

Crusoe

Crusoe

San Francisco, CA
Staff Software Engineer
$215k+/yrOn-site7+ YOEDevOps / SRE

Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer - Site Experience
$217k+/yrOn-site8+ YOEDevOps / SRE

Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer, Ads
$217k+/yrRemote8+ YOEDevOps / SRE

Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.

Reddit

Reddit

United States

Staff Software Engineer, Observability
$217k+/yrRemote7+ YOEDevOps / SRE

Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.

Airbnb

Airbnb

United States

Staff Software Engineer, Service Tools
$212k+/yrRemote9+ YOEDevOps / SRE

Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.