Staff Infrastructure Engineer
Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.
About the job
Responsibilities
- Architect and own the cloud platform used by engineers to deploy services.
- Improve deployment architecture and blast-radius containment through per-service deploy isolation and functional-area slices.
- Own the ECS and EKS footprint, evaluate broader EKS adoption for AI workloads, and design inter-service network connectivity.
- Engineer capacity and autoscaling for spiky workloads, including proactive capacity floors and early drift detection.
- Build a Terraform self-service infrastructure platform with guardrails for standard engineering changes.
- Establish per-team cost attribution and controls across AWS, Datadog, and LLM spending.
- Own Python runtime performance and dependency health, including garbage collection, event-loop contention, runtime limits, framework upgrades, and package upgrades.
- Drive technical decisions across teams and improve engineering practices through architecture reviews, runbooks, and paved-road tooling.
Requirements
- 8+ years in platform, infrastructure, or SRE roles at companies with significant production traffic.
- Deep AWS expertise and production ownership of compute and networking at scale, including ECS, EKS, RDS, networking, and IAM.
- Strong infrastructure-as-code experience, particularly Terraform, including self-service platforms for engineering teams.
- Hands-on autoscaling and capacity engineering experience.
- Container orchestration experience with ECS and/or EKS.
- Track record of making deployments safe and self-service for other teams.
- Staff-level influence, with the ability to drive cross-team decisions without management authority.
Nice to Have
- FinOps and cloud cost optimization experience.
- Kubernetes and EKS depth.
- Observability tooling at scale, particularly Datadog.
- Experience in healthcare or other regulated environments.
- Experience with event-driven systems.
Skills
AWS, ECS, EKS, Rds, Terraform, Kubernetes, IAM, Datadog, Python, Autoscaling, Capacity Engineering, Networking, Infrastructure As Code, Observability, Event-Driven Systems
Similar jobs
DevOps / SRE jobsStaff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.
Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.
Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.