Platform Engineer - Compute Capacity
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
About the job
Responsibilities
- Help maintain the compute capacity plan across regions and instance families.
- Model headroom targets and cost tradeoffs for buffer policy.
- Build automation for reservation acquisition, top-ups, fleet reconciliation, and drift detection.
- Extend infrastructure as code for capacity-related provisioning.
- Develop metrics for saturation, reservation coverage, idle buffer, forecast error, and provisioning latency.
- Build and tune alerts for headroom, quota, and reservation issues.
- Support intake of customer commitments, launches, migrations, and new regions into capacity planning.
- Contribute to right-sizing, instance-family migration, autoscaling, and workload-consolidation initiatives.
- Reduce compute cost per database through right-sizing, placement, and commitment coverage.
- Participate in capacity incident response and post-incident follow-through.
Requirements
- 5+ years of experience in infrastructure engineering, SRE, platform engineering, or capacity engineering, ideally at SaaS or cloud infrastructure scale.
- Production software engineering experience owning and operating services.
- Experience with modern programming languages such as TypeScript, Python, and Go.
- Experience with Infrastructure as Code, including Pulumi or Terraform.
- Experience building or maintaining observability, metrics, and trusted alerting.
- Working knowledge of AWS EC2 instance families and generations, purchase options, capacity reservations, and quota mechanics.
- Financial fluency regarding reservation economics, coverage, utilization, and cost tradeoffs.
- Clear communication across engineering and finance stakeholders.
Nice to Have
- Experience with Prometheus, Grafana, CloudWatch, or similar observability tools.
Compensation and Benefits
- Fully remote work with global hiring.
- WeWork membership or coworking allowance.
- Employee stock ownership plan (ESOP).
- Tech allowance for equipment and work setup.
- Health insurance coverage for employees and dependents.
- Annual company off-sites.
- Flexible, asynchronous work.
- Annual professional development allowance.
Skills
TypeScript, Python, Go, Pulumi, Terraform, Prometheus, Grafana, Amazon Cloudwatch, Aws Ec2, Infrastructure As Code, Capacity Planning, Observability, Autoscaling
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.