Lead Infrastructure Engineer
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.
About the job
Responsibilities
- Own availability, disaster recovery, backup commitments, multi-region failover architecture, and recovery exercises.
- Establish production monitoring, alerting, SLOs, on-call rotation, and incident response processes.
- Own technical controls for SOC 2 and HIPAA, including AWS organization guardrails, least-privilege IAM, KMS/encryption, vulnerability remediation, and continuous audit evidence through Vanta.
- Build CI/CD pipelines, Terraform/IaC foundations, preview environments, and test infrastructure.
- Create guardrails for safe AI coding-agent deployments, including policy-as-code, deploy verification, and agent-operated operations tooling.
- Set platform-work strategy, priorities, status reporting, and operating cadence.
- Architect AI SRE agents, right-size reliability practices, automate compliance controls and audit evidence, and support defined RTO/RPO targets with immutable, restore-tested backups.
Requirements
- Deep AWS experience owning production cloud infrastructure, including IAM, networking, KMS, containers, and managed databases.
- Strong Terraform/IaC and CI/CD expertise.
- Hands-on experience completing at least one SOC 2, HITRUST, or ISO 27001 audit cycle and implementing its technical controls.
- SRE fundamentals covering SLOs, incident management, and disaster recovery design.
- Experience leading engineering teams as a technical lead or engineering manager.
- Ability to break down ambiguous goals, delegate to humans or AI agents, and communicate with non-technical stakeholders.
- Customer-focused mindset and interest in healthcare technology.
Nice-to-haves
- Healthcare or other regulated-industry experience, including HIPAA fluency.
- Policy-as-code experience with OPA or Kyverno.
- Compliance automation experience with Vanta or Drata.
- Experience as a first infrastructure hire or founding/leading a platform team.
- Internal tooling or infrastructure experience for LLM or agent systems.
Compensation and Benefits
- Salary: $200,000–$275,000 annually.
- Hybrid arrangement with 3 days per week in the San Francisco office.
- Unlimited vacation, paid parental leave, medical, dental, and vision insurance.
- Pre-tax commuter benefits, 401(k), significant equity, mentorship, company equipment, home-office setup, and team events/offsites.
Skills
AWS, Terraform, CI/CD, IAM, Kms, Docker, SRE, SLOs, Disaster Recovery, Policy-As-Code, Opa, Kyverno, HIPAA, SOC 2, Python
Similar jobs
DevOps / SRE jobsBuild and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.
Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.