# Release Engineer

**Company:** [Supabase](https://hotfix.jobs/companies/supabase)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** SRE, Kubernetes, AWS, Terraform, Pulumi, Prometheus, Grafana, alertmanager, incident.io, pagerduty, opsgenie, dora metrics, slos, slas
**Posted:** 2026-07-27

> Own reliability, SLOs, and observability for Supabase's deployment pipelines, control plane, and release systems as part of the Release Engineering / SRE team. Drive safe, observable deploys, disaster recovery, incident response, and toil reduction in a fully remote, async environment.

## Job Description

## What You'll Be Responsible For
- Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets.
- Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows.
- Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies).
- Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do.
- Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load.
- Participate in on-call, lead blameless postmortems, and turn findings into runbooks, alerting, and automation that remove toil.
- Improve deployment observability and auditability — a clear record of what shipped where, when, and by whom.
- Document operational procedures — break-glass paths, access models, and runbooks — so reliability knowledge isn't tribal.

## Reliability & Operations
- Define and track SLAs, SLOs, error budgets, and DORA delivery metrics — with meaningful alerting over noise.
- Ensure deployments fail fast and safely when health checks degrade.
- Harden access and break-glass workflows (e.g. scoped self-service) so the right people can act in an incident without unsafe workarounds.
- Partner with product engineering and platform teams to align release practices with reliability and availability targets.

## You Might Be a Good Fit If You
- Have 5+ years in SRE, production operations, platform engineering, or release engineering.
- Have operated production systems at scale and carried on-call for them.
- Are fluent in SLAs, SLOs, error budgets, DORA metrics, and operational KPIs — and the observability tooling behind them (Prometheus, Grafana, Alertmanager, or similar).
- Have led incident response with tooling like incident.io (or PagerDuty / Opsgenie), run blameless postmortems, and driven down MTTD/MTTRO.
- Operate confidently on AWS (multiple accounts, IAM, VPC) in production.
- Are comfortable with infrastructure-as-code (Pulumi, Terraform) and Kubernetes.
- Script and automate to eliminate toil rather than absorb it.
- Communicate clearly with both infrastructure specialists and product engineers.
- Thrive in async, globally distributed teams.
- Are comfortable navigating ambiguity and iterating toward better systems over time.

## What We Offer
- **Fully Remote**: We hire globally. We believe you can do your best work from anywhere. There are no Supabase offices, but we provide a WeWork membership or co-working allowance you can use anywhere in the world.
- **ESOP**: Every team member receives ESOP (equity ownership) in the company. We want everyone to share in the upside of what we’re building together.
- **Tech Allowance**: Use this budget to set up your ideal work environment—laptop, monitor, headphones, or whatever helps you do your best work.
- **Health Benefits**: Supabase covers 100% of health insurance for employees and 80% for dependents, wherever you are. Your wellbeing and your family’s health are important to us.
- **Annual Off-Sites**: Once a year, the entire company gathers in a new city for a week of connection, collaboration, and fun. It’s a highlight of our year.
- **Flexible Work**: We operate asynchronously and trust you to manage your own time. You know what needs to be done and when.
- **Professional Development**: Every team member receives an annual education allowance to spend on learning—courses, books, conferences, or anything that supports your growth.

## Similar roles

- [Operational Technology Engineer](https://hotfix.jobs/jobs/e57e0434-245d-44a8-b378-1c6ab8bf6c2a) - xAI - Memphis, TN
- [Software Engineer III - Data Platform](https://hotfix.jobs/jobs/8ebe115d-a849-434e-a49f-8e5bef3e1f5d) - Idme - Mountain View, CA - $173k – $201k/yr
- [Software Engineer, Networking & Linux Systems](https://hotfix.jobs/jobs/1a5721e0-9d7d-4f3a-8a0f-4562fcb2254a) - Avride - Austin, TX
- [Data Center Engineer](https://hotfix.jobs/jobs/d29f9ff6-7237-4b28-a04c-531a9b90c07d) - Cloudflare - Austin, TX
- [Software Engineer-Platform Engineering](https://hotfix.jobs/jobs/f2500f05-d10d-462d-a2cf-106316c1f090) - Twilio - Remote - $139k – $204k/yr

**Apply:** https://hotfix.jobs/jobs/afc002e9-739a-4adb-b48e-9053a488349b
**Canonical:** https://hotfix.jobs/jobs/afc002e9-739a-4adb-b48e-9053a488349b