Skip to content
HeadwayHeadway

Staff Infrastructure Engineer

Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.

About the job

Responsibilities

  • Architect and own the cloud platform used by engineers to deploy services.
  • Improve deployment architecture and blast-radius containment through per-service deploy isolation and functional-area slices.
  • Own the ECS and EKS footprint, evaluate broader EKS adoption for AI workloads, and design inter-service network connectivity.
  • Engineer capacity and autoscaling for spiky workloads, including proactive capacity floors and early drift detection.
  • Build a Terraform self-service infrastructure platform with guardrails for standard engineering changes.
  • Establish per-team cost attribution and controls across AWS, Datadog, and LLM spending.
  • Own Python runtime performance and dependency health, including garbage collection, event-loop contention, runtime limits, framework upgrades, and package upgrades.
  • Drive technical decisions across teams and improve engineering practices through architecture reviews, runbooks, and paved-road tooling.

Requirements

  • 8+ years in platform, infrastructure, or SRE roles at companies with significant production traffic.
  • Deep AWS expertise and production ownership of compute and networking at scale, including ECS, EKS, RDS, networking, and IAM.
  • Strong infrastructure-as-code experience, particularly Terraform, including self-service platforms for engineering teams.
  • Hands-on autoscaling and capacity engineering experience.
  • Container orchestration experience with ECS and/or EKS.
  • Track record of making deployments safe and self-service for other teams.
  • Staff-level influence, with the ability to drive cross-team decisions without management authority.

Nice to Have

  • FinOps and cloud cost optimization experience.
  • Kubernetes and EKS depth.
  • Observability tooling at scale, particularly Datadog.
  • Experience in healthcare or other regulated environments.
  • Experience with event-driven systems.

Skills

AWS, ECS, EKS, Rds, Terraform, Kubernetes, IAM, Datadog, Python, Autoscaling, Capacity Engineering, Networking, Infrastructure As Code, Observability, Event-Driven Systems

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Deployment Automation Engineer
$250k+/yrOn-site12+ YOEDevOps / SRE

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Software Engineer, DC Infrastructure
$250k+/yrOn-site7+ YOEDevOps / SRE

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.

Datadog

Datadog

Boston, MA
Staff Engineer - Cloud Networks
$244k+/yrHybrid7+ YOEDevOps / SRE

Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.

Skydio

Skydio

San Mateo, CA
Staff Site Reliability Engineer
$240k+/yrRemote8+ YOEDevOps / SRE

Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.