Skip to content

Senior Site Reliability Engineer

Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.

About the job

Responsibilities

  • Own the reliability, performance, and resilience of Garner’s AWS and Kubernetes cloud environments, including AI/ML workloads.
  • Define, measure, and uphold service-level objectives (SLOs) for critical services.
  • Participate in on-call rotations, lead incident response, conduct root-cause analyses, and drive corrective actions through resolution.
  • Build and maintain monitoring, alerting, and observability systems.
  • Translate scaling requirements into automated, composable Terraform infrastructure-as-code deliverables.
  • Optimize cloud costs and performance across compute, storage, and networking.
  • Reduce operational toil and technical debt through AI-assisted automation and monitored, hands-free processes.
  • Establish deployment and observability standards that help engineers ship AI features reliably.
  • Communicate cloud and reliability concepts to technical and non-technical stakeholders.
  • Ensure infrastructure and operations satisfy security and HIPAA compliance obligations.

Requirements

  • 4+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role.
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment; AWS preferred.
  • Experience defining SLOs, building monitoring and alerting, leading incident response, and conducting blameless post-incident reviews.
  • Strong software engineering fundamentals in Python or Go, applied to infrastructure automation.
  • Experience optimizing cloud cost and performance.
  • Fluency with AI tools applied to engineering and operations workflows, or strong motivation to develop this capability.

Nice-to-haves

  • Experience with Kubernetes APIs.
  • Experience supporting AI/ML or data-intensive workloads in production.
  • Experience in security-conscious or regulated environments, including HIPAA or SOC 2.

Technologies

  • AWS
  • Kubernetes
  • Terraform
  • Istio
  • Python
  • Go
  • TypeScript
  • PostgreSQL
  • NATS
  • Datadog
  • GitLab

Compensation

  • Target base compensation: $191,000–$226,000 annually.
  • Eligible for equity incentives and benefits including flexible paid time off, medical/dental/vision plans, 401(k) matching, flexible spending accounts, and Teladoc Health.

Skills

AWS, Kubernetes, Terraform, SLOs, Incident Response, Observability, Datadog, Python, Go, Istio, TypeScript, Postgres, Nats, GitLab

Idme

Idme

McLean, VA
Senior Software Engineer – Platform & Data Infrastructure
$191k+/yrOn-site8+ YOEDevOps / SRE

Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.

Astra

Astra

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Senior Software Engineer - Cloud Infrastructure
$190k+/yrOn-site5+ YOEDevOps / SRE

Build and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.

inKind

inKind

Austin, TX

Senior Platform Engineer
$190k+/yrRemote8+ YOEDevOps / SRE

Own and evolve AWS cloud infrastructure, deployment, reliability, observability, and security for a growing financial and hospitality technology platform. The hands-on role requires 8+ years operating production cloud infrastructure, strong AWS and container orchestration expertise, and experience with migrations and incident response.

Mercury

Mercury

San Francisco, CA
Senior Software Engineer - SRE
$190k+/yrRemote5+ YOEDevOps / SRE

Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.