# Site Reliability Engineer

**Company:** [PostHog](https://hotfix.jobs/companies/posthog)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Kubernetes, EKS, AWS, Terraform, terragrunt, Linux, Argo CD, GitOps, GitHub Actions, cilium, karpenter
**Posted:** 2026-07-16

> Site Reliability Engineer responsible for operating and scaling a large multi-region, multi-account AWS + Kubernetes platform. Focus on automation, IaC with Terraform/Terragrunt, reducing operational toil, and owning production stateful systems end-to-end including on-call.

## Job Description

## What you'll be doing
- Operating EKS clusters across several environments with Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments
- Managing and evolving a multi AWS account organization, including provisioning, networking, access control, and cross-account connectivity
- Maintaining the Terraform/Terragrunt IaC platform - modules, automated plan-on-PR / apply-on-merge pipelines, and safe patterns for shared infrastructure
- Improving operational tooling around deploys, schema changes, backups, restores, and incident response
- Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation
- Optimizing cloud spend as you go
- Participating in on-call and incident response, with a strong focus on making incidents rarer over time

## Requirements
- Deep hands-on experience with Kubernetes in production (EKS preferred). You've debugged node pressure, networking issues, and deployment failures at scale (thousands of nodes)
- Strong experience operating production infrastructure on AWS. Not just one account, but understanding organizational boundaries, IAM, and networking between many
- Experience automating infrastructure using Terraform or Terragrunt at scale, including module design and state management
- Solid understanding of Linux systems (disk, memory, networking, failure modes)
- Experience supporting stateful systems (databases, queues, storage systems, etc.)
- Ability to debug and reason about performance and reliability issues in production
- You're comfortable owning systems end-to-end, including on-call responsibilities

## Nice to have
- Experience with GitOps workflows (ArgoCD) and CI/CD pipelines (GitHub Actions)
- Experience with building AI agent-enabled base-level infra services for teams that move fast
- Familiarity with multi-region infrastructure and the consistency/availability tradeoffs that come with it

## Similar roles

- [Infrastructure Engineer](https://hotfix.jobs/jobs/37da0e47-448c-4ecb-97d9-29d8908ea78c) - Elicit - Oakland, CA
- [Simulation Environments Engineer](https://hotfix.jobs/jobs/3e59b95c-d9a1-4791-a858-85e1cf3abe43) - OpenAI - San Francisco, CA - $230k – $385k/yr
- [Operations Engineer, BizTech](https://hotfix.jobs/jobs/b0de6d38-f928-4f7b-9493-229fd1c6ab62) - Airbnb - Remote - $136k – $160k/yr
- [Systems Integration Engineer, Build Systems | Consumer Devices](https://hotfix.jobs/jobs/4504a8ef-1c33-4742-b405-0358b1508d76) - OpenAI - San Francisco, CA - $293k – $325k/yr
- [Software Engineer, CI Platform Infrastructure](https://hotfix.jobs/jobs/fdfecbed-6550-4136-926b-f191623c3751) - Airbnb - Remote - $162k – $190k/yr

**Apply:** https://hotfix.jobs/jobs/943ece72-acae-43b2-af5e-e37604d0cadc
**Canonical:** https://hotfix.jobs/jobs/943ece72-acae-43b2-af5e-e37604d0cadc