# Staff Infrastructure Engineer

**Company:** [Headway](https://hotfix.jobs/companies/headway)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $265k – $331k/yr
**Experience:** 8+ years
**Skills:** AWS, ECS, EKS, Rds, Terraform, Kubernetes, IAM, Datadog, Python, Autoscaling, Capacity Engineering, Networking, Infrastructure As Code, Observability, Event-Driven Systems
**Posted:** 2026-08-12

> Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.

## Job Description

## Responsibilities
- Architect and own the cloud platform used by engineers to deploy services.
- Improve deployment architecture and blast-radius containment through per-service deploy isolation and functional-area slices.
- Own the ECS and EKS footprint, evaluate broader EKS adoption for AI workloads, and design inter-service network connectivity.
- Engineer capacity and autoscaling for spiky workloads, including proactive capacity floors and early drift detection.
- Build a Terraform self-service infrastructure platform with guardrails for standard engineering changes.
- Establish per-team cost attribution and controls across AWS, Datadog, and LLM spending.
- Own Python runtime performance and dependency health, including garbage collection, event-loop contention, runtime limits, framework upgrades, and package upgrades.
- Drive technical decisions across teams and improve engineering practices through architecture reviews, runbooks, and paved-road tooling.

## Requirements
- **8+ years** in platform, infrastructure, or SRE roles at companies with significant production traffic.
- Deep AWS expertise and production ownership of compute and networking at scale, including ECS, EKS, RDS, networking, and IAM.
- Strong infrastructure-as-code experience, particularly Terraform, including self-service platforms for engineering teams.
- Hands-on autoscaling and capacity engineering experience.
- Container orchestration experience with ECS and/or EKS.
- Track record of making deployments safe and self-service for other teams.
- Staff-level influence, with the ability to drive cross-team decisions without management authority.

## Nice to Have
- FinOps and cloud cost optimization experience.
- Kubernetes and EKS depth.
- Observability tooling at scale, particularly Datadog.
- Experience in healthcare or other regulated environments.
- Experience with event-driven systems.

## Similar jobs

- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/6e086905-4170-407a-a7a4-2e05df0701c5) - Polymarket - New York, NY - $250k – $500k/yr
- [Senior Staff Deployment Automation Engineer](https://hotfix.jobs/jobs/e7c5a05a-7667-45ba-8f89-00349f1e9aaa) - Crusoe - San Francisco, CA - $250k – $300k/yr
- [Senior Staff Software Engineer, DC Infrastructure](https://hotfix.jobs/jobs/5e51b5cc-872b-4556-a065-f5f95a363fad) - Crusoe - San Francisco, CA - $250k – $300k/yr
- [Staff Engineer - Cloud Networks](https://hotfix.jobs/jobs/66aff9c2-57a3-4f6b-bc39-51a61a03c8ae) - Datadog - Boston, MA - $244k – $305k/yr
- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/75edbb52-9f2f-49c0-b1f8-8b59e87fffc3) - Skydio - Remote - $240k – $300k/yr

**Apply:** https://hotfix.jobs/jobs/b1ad3336-2ce1-4ede-a90f-901a178fe108
**Canonical:** https://hotfix.jobs/jobs/b1ad3336-2ce1-4ede-a90f-901a178fe108