Skip to content

Lead Infrastructure Engineer

Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.

About the job

Responsibilities

  • Own availability, disaster recovery, backup commitments, multi-region failover architecture, and recovery exercises.
  • Establish production monitoring, alerting, SLOs, on-call rotation, and incident response processes.
  • Own technical controls for SOC 2 and HIPAA, including AWS organization guardrails, least-privilege IAM, KMS/encryption, vulnerability remediation, and continuous audit evidence through Vanta.
  • Build CI/CD pipelines, Terraform/IaC foundations, preview environments, and test infrastructure.
  • Create guardrails for safe AI coding-agent deployments, including policy-as-code, deploy verification, and agent-operated operations tooling.
  • Set platform-work strategy, priorities, status reporting, and operating cadence.
  • Architect AI SRE agents, right-size reliability practices, automate compliance controls and audit evidence, and support defined RTO/RPO targets with immutable, restore-tested backups.

Requirements

  • Deep AWS experience owning production cloud infrastructure, including IAM, networking, KMS, containers, and managed databases.
  • Strong Terraform/IaC and CI/CD expertise.
  • Hands-on experience completing at least one SOC 2, HITRUST, or ISO 27001 audit cycle and implementing its technical controls.
  • SRE fundamentals covering SLOs, incident management, and disaster recovery design.
  • Experience leading engineering teams as a technical lead or engineering manager.
  • Ability to break down ambiguous goals, delegate to humans or AI agents, and communicate with non-technical stakeholders.
  • Customer-focused mindset and interest in healthcare technology.

Nice-to-haves

  • Healthcare or other regulated-industry experience, including HIPAA fluency.
  • Policy-as-code experience with OPA or Kyverno.
  • Compliance automation experience with Vanta or Drata.
  • Experience as a first infrastructure hire or founding/leading a platform team.
  • Internal tooling or infrastructure experience for LLM or agent systems.

Compensation and Benefits

  • Salary: $200,000–$275,000 annually.
  • Hybrid arrangement with 3 days per week in the San Francisco office.
  • Unlimited vacation, paid parental leave, medical, dental, and vision insurance.
  • Pre-tax commuter benefits, 401(k), significant equity, mentorship, company equipment, home-office setup, and team events/offsites.

Skills

AWS, Terraform, CI/CD, IAM, Kms, Docker, SRE, SLOs, Disaster Recovery, Policy-As-Code, Opa, Kyverno, HIPAA, SOC 2, Python

Skydio

Skydio

San Mateo, CA

Senior Software Engineer, Developer Productivity
$200k+/yrOn-site5+ YOEDevOps / SRE

Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.

Anyscale

Anyscale

San Francisco, CA

Senior Site Reliability Engineer, Platform Infrastructure
$200k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Garner Health

Garner Health

United States

Senior Site Reliability Engineer
$191k+/yrRemote5+ YOEDevOps / SRE

Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.

Idme

Idme

McLean, VA
Senior Software Engineer – Platform & Data Infrastructure
$191k+/yrOn-site8+ YOEDevOps / SRE

Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.