Skip to content
PostHogPostHog

Site Reliability Engineer

Site Reliability Engineer responsible for operating and scaling a large multi-region, multi-account AWS + Kubernetes platform. Focus on automation, IaC with Terraform/Terragrunt, reducing operational toil, and owning production stateful systems end-to-end including on-call.

About the job

What you'll be doing

  • Operating EKS clusters across several environments with Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments
  • Managing and evolving a multi AWS account organization, including provisioning, networking, access control, and cross-account connectivity
  • Maintaining the Terraform/Terragrunt IaC platform - modules, automated plan-on-PR / apply-on-merge pipelines, and safe patterns for shared infrastructure
  • Improving operational tooling around deploys, schema changes, backups, restores, and incident response
  • Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation
  • Optimizing cloud spend as you go
  • Participating in on-call and incident response, with a strong focus on making incidents rarer over time

Requirements

  • Deep hands-on experience with Kubernetes in production (EKS preferred). You've debugged node pressure, networking issues, and deployment failures at scale (thousands of nodes)
  • Strong experience operating production infrastructure on AWS. Not just one account, but understanding organizational boundaries, IAM, and networking between many
  • Experience automating infrastructure using Terraform or Terragrunt at scale, including module design and state management
  • Solid understanding of Linux systems (disk, memory, networking, failure modes)
  • Experience supporting stateful systems (databases, queues, storage systems, etc.)
  • Ability to debug and reason about performance and reliability issues in production
  • You're comfortable owning systems end-to-end, including on-call responsibilities

Nice to have

  • Experience with GitOps workflows (ArgoCD) and CI/CD pipelines (GitHub Actions)
  • Experience with building AI agent-enabled base-level infra services for teams that move fast
  • Familiarity with multi-region infrastructure and the consistency/availability tradeoffs that come with it

Skills

Kubernetes, EKS, AWS, Terraform, Terragrunt, Linux, Argo CD, GitOps, GitHub Actions, Cilium, Karpenter

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.