# Platform Engineer - Compute Capacity

**Company:** [Supabase](https://hotfix.jobs/companies/supabase)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** TypeScript, Python, Go, Pulumi, Terraform, Prometheus, Grafana, Amazon Cloudwatch, Aws Ec2, Infrastructure As Code, Capacity Planning, Observability, Autoscaling
**Posted:** 2026-09-02

> Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

## Job Description

## Responsibilities
- Help maintain the compute capacity plan across regions and instance families.
- Model headroom targets and cost tradeoffs for buffer policy.
- Build automation for reservation acquisition, top-ups, fleet reconciliation, and drift detection.
- Extend infrastructure as code for capacity-related provisioning.
- Develop metrics for saturation, reservation coverage, idle buffer, forecast error, and provisioning latency.
- Build and tune alerts for headroom, quota, and reservation issues.
- Support intake of customer commitments, launches, migrations, and new regions into capacity planning.
- Contribute to right-sizing, instance-family migration, autoscaling, and workload-consolidation initiatives.
- Reduce compute cost per database through right-sizing, placement, and commitment coverage.
- Participate in capacity incident response and post-incident follow-through.

## Requirements
- 5+ years of experience in infrastructure engineering, SRE, platform engineering, or capacity engineering, ideally at SaaS or cloud infrastructure scale.
- Production software engineering experience owning and operating services.
- Experience with modern programming languages such as TypeScript, Python, and Go.
- Experience with Infrastructure as Code, including Pulumi or Terraform.
- Experience building or maintaining observability, metrics, and trusted alerting.
- Working knowledge of AWS EC2 instance families and generations, purchase options, capacity reservations, and quota mechanics.
- Financial fluency regarding reservation economics, coverage, utilization, and cost tradeoffs.
- Clear communication across engineering and finance stakeholders.

## Nice to Have
- Experience with Prometheus, Grafana, CloudWatch, or similar observability tools.

## Compensation and Benefits
- Fully remote work with global hiring.
- WeWork membership or coworking allowance.
- Employee stock ownership plan (ESOP).
- Tech allowance for equipment and work setup.
- Health insurance coverage for employees and dependents.
- Annual company off-sites.
- Flexible, asynchronous work.
- Annual professional development allowance.

## Similar jobs

- [Software Engineer: Resiliency - Deploy at Scale](https://hotfix.jobs/jobs/21bb6f4f-daea-46ab-bac8-fff083d17451) - Cloudflare - London, United Kingdom
- [Release Engineer - Data Plane Internal Tooling and Productivity](https://hotfix.jobs/jobs/ee184847-0b76-42f4-8c1f-157f199626d3) - Clickhouse - Remote
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Capacity Ops Engineer](https://hotfix.jobs/jobs/f1904714-7dd3-4ee3-9e7a-e4fcf52083bd) - Baseten - San Francisco, CA - $170k – $230k/yr
- [IT Security and Automation Engineer](https://hotfix.jobs/jobs/604b87b5-13a2-4bba-88b2-f7d0fbbad141) - Teleport - Remote - $149k – $258k/yr

**Apply:** https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4
**Canonical:** https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4