Skip to content
TabsTabs

Staff Site Reliability Engineer

Leads infrastructure evolution, CI/CD, observability, and reliability for scaling platform. Requires 10+ years SRE/infra experience, AWS expertise, and systems thinking. Onsite in NYC.

About the job

What You’ll Own

  • AWS infrastructure direction and platform evolution, including migration from ECS/Fargate toward modern, scalable runtime
  • CI/CD systems emphasizing developer experience, safety, and automation (GitHub Actions; maturing CD)
  • Ephemeral environments and preview deploys
  • Observability standards (metrics, logs, tracing, alert hygiene, dashboards, SLO development)
  • Incident response, postmortems, and reliability culture

What You’ll Do

  • Define and evolve reliability standards, SLIs, SLOs, and error budgets
  • Improve observability, alerting, and incident processes
  • Lead high-severity incidents and drive follow-ups
  • Partner with teams to design resilient, scalable systems
  • Build automation to reduce toil and risk
  • Mentor engineers and influence best practices

Who You Are

  • Run production systems on AWS and lead platform change
  • Think in systems: risk, rollback, blast radius, feedback loops
  • Treat CI/CD and environments as self-serve products
  • Influence through trust and clarity
  • Balance pragmatism with system health
  • Value learning from failure
  • Communicate clearly across teams

Experience

  • 10+ years in SRE, infrastructure, or backend engineering
  • Strong software engineering in modern languages
  • Expertise in distributed systems at scale
  • Deep AWS, observability tooling, CI/CD experience
  • Comfortable with ambiguity

Perks and Benefits

  • Competitive compensation and equity
  • Unlimited PTO
  • Up to 100% employer-covered healthcare
  • Meals provided
  • Parental leave, commuter benefits, 401k

Skills

AWS, Kubernetes, GitHub Actions, CI/CD, Slo, Sli, Observability, Distributed Systems, Incident Response, Terraform

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Operations
$195k+/yrOn-site8+ YOEDevOps / SRE

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.

Shield AI

Shield AI

San Diego, CA

Senior Staff Lead Site Reliability Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.

Airbnb

Airbnb

United States

Staff Software Engineer, Service Tools
$212k+/yrRemote9+ YOEDevOps / SRE

Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.

Temporal

Temporal

United States

Staff Software Engineer, Traffic
$212k+/yrRemote8+ YOEDevOps / SRE

Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.