Skip to content
HarveyHarvey

Staff Software Engineer, Site Reliability Engineer (SRE)

Staff SRE ensures reliability, scalability, and performance of legal AI platform across global regions. Leads incident management, automates operations, optimizes infrastructure costs, and mentors teams. Requires 10+ years SRE experience, IaC, cloud platforms, and observability tools.

About the job

What You’ll Do

  • Design, implement, and manage monitoring, alerting, and infrastructure resources (compute, storage, networking) across 50+ global regions
  • Lead incident management processes, including postmortems, root cause analyses, and driving actionable improvements
  • Automate operational tasks and workflows, building tools and processes for capacity planning, graceful rollouts, and safe data access to maintain high reliability and reduce manual intervention
  • Establish best practices for security, compliance, and reliability and collaborate across teams to drive these principles throughout the software lifecycle
  • Optimize infrastructure costs through strategic capacity planning and build-versus-buy decisions while maintaining system performance, reliability, and functionality
  • Provide technical mentorship and leadership, promoting best practices and fostering team growth

What You Have

  • 10+ years of experience in Site Reliability Engineering or similar roles supporting production environments, with proven ability to mentor and guide technical teams
  • Expertise in infrastructure as code (IaC) tools (Pulumi, Terraform, CloudFormation, etc.)
  • Deep familiarity with observability tools (Datadog, Sentry, etc.) and incident response practices (PagerDuty, IncidentIO, etc.)
  • Proficiency with cloud infrastructure platforms (Azure, GCP, AWS, etc.)
  • Strong programming skills (Python, Bash, Go, or similar languages)
  • Proven track record of diagnosing complex system problems and implementing durable solutions
  • Solid understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles
  • Excellent problem-solving skills, meticulous attention to detail, and a commitment to operational excellence

Compensation

$238,000 - $290,000 USD

Skills

Terraform, Pulumi, Kubernetes, Datadog, Pagerduty, Python, Go, AWS, GCP, Azure, CI/CD

Skydio

Skydio

San Mateo, CA
Staff Site Reliability Engineer
$240k+/yrRemote8+ YOEDevOps / SRE

Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.

Datadog

Datadog

Boston, MA
Staff Engineer - Cloud Networks
$244k+/yrHybrid7+ YOEDevOps / SRE

Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Deployment Automation Engineer
$250k+/yrOn-site12+ YOEDevOps / SRE

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Software Engineer, DC Infrastructure
$250k+/yrOn-site7+ YOEDevOps / SRE

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.