Staff Infrastructure Engineer
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
About the job
Responsibilities
- Own the architecture, operating model, and modernization roadmap for cloud infrastructure and shared services, including systems without clear ownership.
- Design, build, and operate AWS and Kubernetes infrastructure as code using Terraform, delivered through GitOps with ArgoCD and CI/CD with GitHub Actions.
- Establish safe, repeatable standards for AI-assisted infrastructure work and improve team fluency with these practices.
- Strengthen security and compliance through least-privilege IAM and RBAC, identity and access management, network boundaries, and secrets hygiene aligned with SOC2 expectations.
- Optimize infrastructure for reliability, scalability, developer experience, and cost.
- Provide design review, mentorship, documentation, and enablement across the infrastructure team.
- Participate in and improve the shared US-business-hours on-call rotation, incident response, alerting, and runbooks.
Requirements
- 8+ years of infrastructure, cloud, or platform engineering experience, including deep hands-on AWS experience operating production systems at scale.
- Strong Terraform infrastructure-as-code proficiency.
- Experience owning and operating Kubernetes platforms.
- Track record of taking ownership of ambiguous, inherited, or under-documented systems and bringing them to a well-operated state.
- Fluency with AI-assisted engineering tools and sound judgment about their safe application.
- Security and compliance experience in a regulated environment such as HIPAA or SOC2.
- Staff-level technical partnership and influence through standards, reviews, and enablement.
- Demonstrated expertise across multiple infrastructure domains and breadth across adjacent domains.
AI-Assisted Engineering Expectations
- Set standards and safe patterns for using AI assistants and platforms with Terraform and script authoring, runbook and documentation generation, log and query analysis, and support triage.
- Ensure all AI-generated output is reviewed and validated before production use.
Nice-to-Haves
- Okta, OIDC, or SAML.
- Envoy or similar API gateway and service networking technologies.
- Python, Go, or Bash scripting and automation.
- FinOps or cloud cost-optimization experience.
- Platform-scale observability covering metrics, logs, tracing, and alerting.
- Snowflake or similar cloud data warehouses.
Compensation and Benefits
- Annual base pay range: $187,000–$265,000 USD.
- Potential performance-based bonus and equity awards.
- Health, dental, and vision insurance.
- Flexible time off and holidays.
- 401(k) with company match.
- Disability and life insurance.
- Leaves of absence in accordance with applicable laws and company policy.
Skills
AWS, Kubernetes, Terraform, Argo CD, GitHub Actions, IAM, RBAC, HIPAA, Soc2, Python, Go, Bash, Okta, Envoy, Snowflake
Similar jobs
DevOps / SRE jobsLeads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Staff DevSecOps Engineer designing and automating security controls across AWS infrastructure, containers, CI/CD, and platform services. Requires 7+ years of related experience plus expertise in cloud security, infrastructure as code, hardened images, vulnerability scanning, identity, and secrets management.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.