Skip to content
SupabaseSupabase

Site Reliability Engineer

SRE embedded in Service Operations to establish reliability practices, frameworks, and feedback loops across engineering teams. Focus on SLOs/SLIs, ORR processes, incident-to-improvement pipelines, and influencing without authority in a distributed environment.

About the job

What You'll Own

  • Partner with service teams to define meaningful SLIs and SLOs grounded in customer experience, and build the error budget policies that turn them into engineering decisions
  • Own and evolve the Operational Readiness Review (ORR) process — conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation
  • Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes
  • Act as the reliability expert teams pull in for architecture reviews, failure mode analysis, dependency mapping, and resilience design
  • Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it
  • Help teams design sustainable on-call practices: alert quality, escalation paths, runbook coverage, and noise reduction
  • Track and report on org-wide operational maturity, surfacing systemic gaps and driving remediation

Requirements

  • 7+ years of experience in SRE, production engineering, or reliability-focused roles, including experience shaping SRE practices and driving adoption across engineering teams
  • Software engineering mindset — write code and build tools, not just configure them
  • Hands-on experience defining and operationalizing SLOs/SLIs at scale, including error budget policies that actually influenced engineering decisions
  • Deep experience with incident response, postmortem facilitation, and turning incident learnings into systemic improvements
  • Worked with large-scale multi-tenant systems (bonus: managed database platforms or Postgres)
  • Proficient with cloud infrastructure (AWS preferred) and infrastructure-as-code (Pulumi preferred, Terraform/CDK also acceptable)
  • Communicate clearly and persuasively — this role requires influencing without authority across a distributed org
  • Experience in async or globally distributed teams
  • Energized by making other teams more effective rather than being the one who fixes everything

Nice to Have

  • Experience with Kubernetes-based platform operations
  • Familiarity with OpenTelemetry, VictoriaMetrics, Grafana, or similar observability tooling
  • Experience building developer-facing reliability tooling (SLO dashboards, ORR frameworks, toil tracking, DORA metrics)

Skills

SRE, SLOs, Slis, Error Budgets, Incident Response, Postmortems, AWS, Pulumi, Terraform, Kubernetes, OpenTelemetry, Grafana, Observability

Shield AI

Shield AI

San Diego, CA
Senior Platform Engineer
$141k+/yrHybrid7+ YOEDevOps / SRE

Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.

Shield AI

Shield AI

San Mateo, CA
Senior Network Engineer
$140k+/yrOn-site6+ YOEDevOps / SRE

Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.

Astra

Astra

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
CA$95k+/yrRemote5+ YOEDevOps / SRE

Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
No salary listedRemote5+ YOEDevOps / SRE

Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.