Skip to content
HarveyHarvey

Senior Software Engineer, Site Reliability Engineer (SRE)

Senior SRE ensures reliability, scalability, and performance of legal AI platform by managing global infrastructure, leading incident response, automating operations, and optimizing costs. Requires 5+ years SRE experience, IaC expertise, cloud proficiency, and strong programming skills.

About the job

What You’ll Do

  • Design, implement, and manage monitoring, alerting, and infrastructure resources (compute, storage, networking) across 50+ global regions
  • Lead incident management processes, including postmortems, root cause analyses, and driving actionable improvements
  • Automate operational tasks and workflows, building tools and processes for capacity planning, graceful rollouts, and safe data access to maintain high reliability and reduce manual intervention
  • Collaborate across teams to drive reliability, security, and compliance throughout the software lifecycle
  • Optimize infrastructure costs through strategic capacity planning and build-versus-buy decisions while maintaining system performance, reliability, and functionality

What You Have

  • 5+ years of experience in Site Reliability Engineering or similar roles supporting production environments
  • Expertise in infrastructure as code(IaC) tools (Pulumi, Terraform, CloudFormation, etc.)
  • Deep familiarity with observability tools (Datadog, Sentry, etc.) and incident response practices (PagerDuty, IncidentIO, etc.)
  • Proficiency with cloud infrastructure platforms (Azure, GCP, AWS, etc.)
  • Strong programming skills (Python, Bash, Go, or similar languages)
  • Proven track record of diagnosing complex system problems and implementing durable solutions
  • Solid understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles
  • Excellent problem-solving skills, meticulous attention to detail, and a commitment to operational excellence

Compensation Range

$200,000 - $260,000 USD

Skills

Terraform, Pulumi, Kubernetes, Datadog, Sentry, Pagerduty, AWS, GCP, Azure, Python, Go, Bash, CI/CD, CloudFormation

Skydio

Skydio

San Mateo, CA

Senior Software Engineer, Developer Productivity
$200k+/yrOn-site5+ YOEDevOps / SRE

Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.

Anyscale

Anyscale

San Francisco, CA

Senior Site Reliability Engineer, Platform Infrastructure
$200k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.

Onos Health

Onos Health

San Francisco, CA

Lead Infrastructure Engineer
$200k+/yrHybrid7+ YOEDevOps / SRE

Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Garner Health

Garner Health

United States

Senior Site Reliability Engineer
$191k+/yrRemote5+ YOEDevOps / SRE

Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.