Skip to content
GleanGlean

Lead Site Reliability Engineer

Leads SRE team to ensure high availability, scalability, and reliability of cloud services through automation, incident management, and technical leadership. Requires 8+ years SRE experience, team management, and expertise in cloud platforms and containerization.

About the job

Responsibilities

  • Technical Leadership and Mentorship: Drive technical excellence, set best practices for incident management, performance optimization, and automation. Influence architectural decisions and ensure high-quality systems.
  • Ensure high availability: Implement resilient cloud architectures and proactively resolve bottlenecks.
  • Incident Management: Participate in oncall rotation, foster blameless postmortems, and optimize on-call processes.
  • Automation and Tooling: Develop scripts and tools to streamline deployment, monitoring, and management.
  • Performance Optimization: Optimize infrastructure for performance, scalability, and cost-effectiveness.
  • Security and Compliance: Collaborate on security best practices.
  • Monitoring and Alerting: Design monitoring systems, dashboards, and playbooks.
  • Software Development Consultation: Participate in design reviews and provide SRE insights.

Requirements

  • Bachelor’s degree in Computer Science or equivalent.
  • 8+ years in senior SRE or similar role managing cloud services.
  • 5+ years software development experience.
  • 3+ years managing teams and distributed systems in cloud.
  • Strong knowledge of GCP, AWS, or Azure.
  • Experience with Docker, Kubernetes, Terraform.
  • Understanding of networking, security, and SRE practices.
  • Proficiency in monitoring and alerting tools.

Compensation & Benefits

Base salary: $200,000 - $260,000 annually. Comprehensive benefits including medical, dental, vision, 401k, stipends, and company events.

Skills

Kubernetes, Docker, Terraform, GCP, AWS, Azure, Monitoring Tools, Alerting Tools, Infrastructure As Code, Distributed Systems

Skydio

Skydio

San Mateo, CA

Senior Software Engineer, Developer Productivity
$200k+/yrOn-site5+ YOEDevOps / SRE

Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.

Anyscale

Anyscale

San Francisco, CA

Senior Site Reliability Engineer, Platform Infrastructure
$200k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.

Onos Health

Onos Health

San Francisco, CA

Lead Infrastructure Engineer
$200k+/yrHybrid7+ YOEDevOps / SRE

Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Garner Health

Garner Health

United States

Senior Site Reliability Engineer
$191k+/yrRemote5+ YOEDevOps / SRE

Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.