Skip to content
Shield AIShield AI

Senior Staff Lead Site Reliability Engineer

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

About the job

Responsibilities

  • Define and implement service-level indicators (SLIs), service-level objectives (SLOs), and other reliability measures.
  • Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services.
  • Lead technical responses to complex incidents and drive root-cause analysis through resolution.
  • Identify recurring failure modes and work with engineering teams to eliminate them.
  • Improve system resilience through automation, testing, capacity planning, and failure recovery.
  • Develop tooling and automation to reduce manual operational work.
  • Partner with product and platform teams to incorporate reliability requirements into system design.
  • Establish incident response practices that improve detection, diagnosis, communication, and recovery.
  • Mentor product engineers and SRE and Cloud Engineering teammates.
  • Define and manage short- and long-term SRE roadmaps and distribute work across teammates.

Requirements

  • 7+ years of experience in SRE, software engineering, infrastructure engineering, or a related field.
  • Experience operating production services with defined availability and reliability requirements.
  • Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices.
  • Experience designing and operating infrastructure in AWS or another major cloud environment.
  • Experience with infrastructure as code and automated infrastructure provisioning.
  • Experience supporting containerized applications and distributed systems.
  • Experience developing operational tooling or automation using Python, Go, or a similar language.
  • Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services.
  • Experience leading incident response and root-cause analysis across engineering teams.
  • Experience leading and executing a team’s technical vision over multi-quarter timelines.

Nice-to-haves

  • Experience establishing or maturing an SRE function within an engineering organization.
  • Experience with Kubernetes and cloud-native observability systems.
  • Experience operating systems in regulated or compliance-driven environments.
  • Experience with capacity planning, performance analysis, and cloud cost management.
  • Experience supporting shared infrastructure across multiple products or engineering organizations.
  • Experience building roadmaps in ticketing systems.
  • Experience mentoring other engineers.

Skills

AWS, Infrastructure As Code, Kubernetes, Python, Go, Monitoring, Alerting, Logging, Distributed Systems, Incident Response, Root-Cause Analysis, Capacity Planning, Observability

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.

Komodo Health

Komodo Health

United States

Staff Infrastructure Engineer
$187k+/yrRemote8+ YOEDevOps / SRE

Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.

VGS

VGS

United States
Senior Staff Infrastructure Engineer
$185k+/yrRemote10+ YOEDevOps / SRE

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Operations
$195k+/yrOn-site8+ YOEDevOps / SRE

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.

Shield AI

Shield AI

San Mateo, CA
Staff DevSecOps Engineer
$182k+/yrOn-site7+ YOEDevOps / SRE

Staff DevSecOps Engineer designing and automating security controls across AWS infrastructure, containers, CI/CD, and platform services. Requires 7+ years of related experience plus expertise in cloud security, infrastructure as code, hardened images, vulnerability scanning, identity, and secrets management.