# Lead Infrastructure Engineer

**Company:** [Onos Health](https://hotfix.jobs/companies/onos-health)
**Location:** San Francisco, CA
**Role:** DevOps / SRE
**Salary:** $200k – $275k/yr
**Experience:** 7+ years
**Skills:** AWS, Terraform, CI/CD, IAM, Kms, Docker, SRE, SLOs, Disaster Recovery, Policy-As-Code, Opa, Kyverno, HIPAA, SOC 2, Python
**Posted:** 2026-08-26

> Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.

## Job Description

## Responsibilities
- Own availability, disaster recovery, backup commitments, multi-region failover architecture, and recovery exercises.
- Establish production monitoring, alerting, SLOs, on-call rotation, and incident response processes.
- Own technical controls for SOC 2 and HIPAA, including AWS organization guardrails, least-privilege IAM, KMS/encryption, vulnerability remediation, and continuous audit evidence through Vanta.
- Build CI/CD pipelines, Terraform/IaC foundations, preview environments, and test infrastructure.
- Create guardrails for safe AI coding-agent deployments, including policy-as-code, deploy verification, and agent-operated operations tooling.
- Set platform-work strategy, priorities, status reporting, and operating cadence.
- Architect AI SRE agents, right-size reliability practices, automate compliance controls and audit evidence, and support defined RTO/RPO targets with immutable, restore-tested backups.

## Requirements
- Deep AWS experience owning production cloud infrastructure, including IAM, networking, KMS, containers, and managed databases.
- Strong Terraform/IaC and CI/CD expertise.
- Hands-on experience completing at least one SOC 2, HITRUST, or ISO 27001 audit cycle and implementing its technical controls.
- SRE fundamentals covering SLOs, incident management, and disaster recovery design.
- Experience leading engineering teams as a technical lead or engineering manager.
- Ability to break down ambiguous goals, delegate to humans or AI agents, and communicate with non-technical stakeholders.
- Customer-focused mindset and interest in healthcare technology.

## Nice-to-haves
- Healthcare or other regulated-industry experience, including HIPAA fluency.
- Policy-as-code experience with OPA or Kyverno.
- Compliance automation experience with Vanta or Drata.
- Experience as a first infrastructure hire or founding/leading a platform team.
- Internal tooling or infrastructure experience for LLM or agent systems.

## Compensation and Benefits
- Salary: $200,000–$275,000 annually.
- Hybrid arrangement with 3 days per week in the San Francisco office.
- Unlimited vacation, paid parental leave, medical, dental, and vision insurance.
- Pre-tax commuter benefits, 401(k), significant equity, mentorship, company equipment, home-office setup, and team events/offsites.

## Similar jobs

- [Senior Software Engineer, Developer Productivity](https://hotfix.jobs/jobs/2862734a-5d2f-4f87-8635-a4c3d042bbe8) - Skydio - San Mateo, CA - $200k – $240k/yr
- [Senior Site Reliability Engineer, Platform Infrastructure](https://hotfix.jobs/jobs/4501ca74-99c6-494f-93a1-0184fd51bfa5) - Anyscale - San Francisco, CA - $200k – $240k/yr
- [Senior Production Engineer](https://hotfix.jobs/jobs/9ee5879e-954d-4681-ad0c-816d7151f874) - Lightspark - Remote - $200k – $238k/yr
- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/c3596786-c3a3-4d93-92ca-c15f6acd5b24) - Garner Health - Remote - $191k – $226k/yr
- [Senior Software Engineer – Platform & Data Infrastructure](https://hotfix.jobs/jobs/22c23d5c-1596-42f5-b2da-6f663d20f3d3) - Idme - McLean, VA - $191k – $214k/yr

**Apply:** https://hotfix.jobs/jobs/a9b6d4e6-46bd-4038-b15b-02f33481c182
**Canonical:** https://hotfix.jobs/jobs/a9b6d4e6-46bd-4038-b15b-02f33481c182