Skip to content
PostHogPostHogUnited States

Site Reliability Engineer

Site Reliability Engineer responsible for operating and scaling a large multi-region, multi-account AWS + Kubernetes platform. Focus on automation, IaC with Terraform/Terragrunt, reducing operational toil, and owning production stateful systems end-to-end including on-call.

Salary not listed
Remote5+ YOEDevOps / SRE

About the role

What you'll be doing

  • Operating EKS clusters across several environments with Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments
  • Managing and evolving a multi AWS account organization, including provisioning, networking, access control, and cross-account connectivity
  • Maintaining the Terraform/Terragrunt IaC platform - modules, automated plan-on-PR / apply-on-merge pipelines, and safe patterns for shared infrastructure
  • Improving operational tooling around deploys, schema changes, backups, restores, and incident response
  • Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation
  • Optimizing cloud spend as you go
  • Participating in on-call and incident response, with a strong focus on making incidents rarer over time

Requirements

  • Deep hands-on experience with Kubernetes in production (EKS preferred). You've debugged node pressure, networking issues, and deployment failures at scale (thousands of nodes)
  • Strong experience operating production infrastructure on AWS. Not just one account, but understanding organizational boundaries, IAM, and networking between many
  • Experience automating infrastructure using Terraform or Terragrunt at scale, including module design and state management
  • Solid understanding of Linux systems (disk, memory, networking, failure modes)
  • Experience supporting stateful systems (databases, queues, storage systems, etc.)
  • Ability to debug and reason about performance and reliability issues in production
  • You're comfortable owning systems end-to-end, including on-call responsibilities

Nice to have

  • Experience with GitOps workflows (ArgoCD) and CI/CD pipelines (GitHub Actions)
  • Experience with building AI agent-enabled base-level infra services for teams that move fast
  • Familiarity with multi-region infrastructure and the consistency/availability tradeoffs that come with it

Skills

KubernetesEKSAWSTerraformterragruntLinuxArgo CDGitOpsGitHub Actionsciliumkarpenter

Similar roles

DevOps / SRE jobs
Elicit

Infrastructure Engineer

ElicitOakland, CA

Own and evolve Elicit's cloud infrastructure platform (AWS/GCP, Kubernetes, Terraform) to support scalable single-tenant enterprise deployments. Build observability, compliance (SOC 2), cost optimization, and developer experience while contributing to backend systems where infra meets application logic. Requires 5+ years infrastructure/SRE experience, strong Terraform and K8s expertise, and enthusiasm for AI coding agents.

Salary not listed
On-site5+ YOEDevOps / SRE
OpenAI

Simulation Environments Engineer

OpenAISan Francisco, CA

Build and maintain CI/CD pipelines, orchestration, and automation for large-scale robotics simulation (SIL/HIL) to support model training, evaluation, and RL workloads at OpenAI. Requires strong infra, distributed systems, and Python/C++/Rust experience.

230k – 385k/yr
Hybrid5+ YOEDevOps / SRE
Airbnb

Operations Engineer, BizTech

AirbnbUnited States

Operations Engineer using AI, LLMs, and intelligent automation to triage tickets, accelerate incident response, build self-healing observability, and automate repetitive operational work in Airbnb's BizTech Global Operations team.

136k – 160k/yr
Remote3+ YOEDevOps / SRE
OpenAI

Systems Integration Engineer, Build Systems | Consumer Devices

OpenAISan Francisco, CA

Build and evolve Bazel, Yocto, and Buildkite-based CI systems for OpenAI consumer device software. Focus on hermetic builds, remote caching, test optimization, observability, and AI-powered failure analysis to accelerate reliable shipping. Requires 5+ years building developer infrastructure at scale.

293k – 325k/yr
Hybrid5+ YOEDevOps / SRE
Airbnb

Software Engineer, CI Platform Infrastructure

AirbnbUnited States

Build and optimize a next-generation CI platform infrastructure for workflow orchestration, scheduling, caching, and autoscaling to accelerate software development for engineers and AI coding agents at scale. Requires interest in distributed systems and knowledge of Kubernetes, EC2, Golang, and Docker.

162k – 190k/yr
RemoteDevOps / SRE