# Senior Site Reliability Engineer

**Company:** [Garner Health](https://hotfix.jobs/companies/garner-health)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $191k – $226k/yr
**Experience:** 5+ years
**Skills:** AWS, Kubernetes, Terraform, SLOs, Incident Response, Observability, Datadog, Python, Go, Istio, TypeScript, Postgres, Nats, GitLab
**Posted:** 2026-09-02

> Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.

## Job Description

## Responsibilities
- Own the reliability, performance, and resilience of Garner’s AWS and Kubernetes cloud environments, including AI/ML workloads.
- Define, measure, and uphold service-level objectives (SLOs) for critical services.
- Participate in on-call rotations, lead incident response, conduct root-cause analyses, and drive corrective actions through resolution.
- Build and maintain monitoring, alerting, and observability systems.
- Translate scaling requirements into automated, composable Terraform infrastructure-as-code deliverables.
- Optimize cloud costs and performance across compute, storage, and networking.
- Reduce operational toil and technical debt through AI-assisted automation and monitored, hands-free processes.
- Establish deployment and observability standards that help engineers ship AI features reliably.
- Communicate cloud and reliability concepts to technical and non-technical stakeholders.
- Ensure infrastructure and operations satisfy security and HIPAA compliance obligations.

## Requirements
- 4+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role.
- Deep expertise with Kubernetes and Terraform in a cloud-first environment; AWS preferred.
- Experience defining SLOs, building monitoring and alerting, leading incident response, and conducting blameless post-incident reviews.
- Strong software engineering fundamentals in Python or Go, applied to infrastructure automation.
- Experience optimizing cloud cost and performance.
- Fluency with AI tools applied to engineering and operations workflows, or strong motivation to develop this capability.

## Nice-to-haves
- Experience with Kubernetes APIs.
- Experience supporting AI/ML or data-intensive workloads in production.
- Experience in security-conscious or regulated environments, including HIPAA or SOC 2.

## Technologies
- AWS
- Kubernetes
- Terraform
- Istio
- Python
- Go
- TypeScript
- PostgreSQL
- NATS
- Datadog
- GitLab

## Compensation
- Target base compensation: $191,000–$226,000 annually.
- Eligible for equity incentives and benefits including flexible paid time off, medical/dental/vision plans, 401(k) matching, flexible spending accounts, and Teladoc Health.

## Similar jobs

- [Senior Software Engineer – Platform & Data Infrastructure](https://hotfix.jobs/jobs/22c23d5c-1596-42f5-b2da-6f663d20f3d3) - Idme - McLean, VA - $191k – $214k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/3317571d-7a67-4059-b23e-af1a9033cf96) - Astra - Remote - $190k – $230k/yr
- [Senior Software Engineer - Cloud Infrastructure](https://hotfix.jobs/jobs/c60e0ab5-199d-42dc-8516-72524bd54fc6) - Applied Intuition - Sunnyvale, CA - $190k – $270k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/aedea6ac-e83e-46e3-adf9-3a43d2d8cf5e) - inKind - Remote - $190k – $200k/yr
- [Senior Software Engineer - SRE](https://hotfix.jobs/jobs/2a0ae39c-973a-4bd9-a509-60717c615ef6) - Mercury - Remote - $190k – $251k/yr

**Apply:** https://hotfix.jobs/jobs/c3596786-c3a3-4d93-92ca-c15f6acd5b24
**Canonical:** https://hotfix.jobs/jobs/c3596786-c3a3-4d93-92ca-c15f6acd5b24