Site Reliability Engineer responsible for operating and scaling a large multi-region, multi-account AWS + Kubernetes platform. Focus on automation, IaC with Terraform/Terragrunt, reducing operational toil, and owning production stateful systems end-to-end including on-call.
Salary not listed
Remote5+ YOEDevOps / SRE
About the role
What you'll be doing
Operating EKS clusters across several environments with Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments
Managing and evolving a multi AWS account organization, including provisioning, networking, access control, and cross-account connectivity
Maintaining the Terraform/Terragrunt IaC platform - modules, automated plan-on-PR / apply-on-merge pipelines, and safe patterns for shared infrastructure
Improving operational tooling around deploys, schema changes, backups, restores, and incident response
Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation
Optimizing cloud spend as you go
Participating in on-call and incident response, with a strong focus on making incidents rarer over time
Requirements
Deep hands-on experience with Kubernetes in production (EKS preferred). You've debugged node pressure, networking issues, and deployment failures at scale (thousands of nodes)
Strong experience operating production infrastructure on AWS. Not just one account, but understanding organizational boundaries, IAM, and networking between many
Experience automating infrastructure using Terraform or Terragrunt at scale, including module design and state management
Solid understanding of Linux systems (disk, memory, networking, failure modes)
Experience supporting stateful systems (databases, queues, storage systems, etc.)
Ability to debug and reason about performance and reliability issues in production
You're comfortable owning systems end-to-end, including on-call responsibilities
Nice to have
Experience with GitOps workflows (ArgoCD) and CI/CD pipelines (GitHub Actions)
Experience with building AI agent-enabled base-level infra services for teams that move fast
Familiarity with multi-region infrastructure and the consistency/availability tradeoffs that come with it
Own and evolve Elicit's cloud infrastructure platform (AWS/GCP, Kubernetes, Terraform) to support scalable single-tenant enterprise deployments. Build observability, compliance (SOC 2), cost optimization, and developer experience while contributing to backend systems where infra meets application logic. Requires 5+ years infrastructure/SRE experience, strong Terraform and K8s expertise, and enthusiasm for AI coding agents.
Salary not listed
On-site5+ YOEDevOps / SRE
Simulation Environments Engineer
OpenAISan Francisco, CA
Build and maintain CI/CD pipelines, orchestration, and automation for large-scale robotics simulation (SIL/HIL) to support model training, evaluation, and RL workloads at OpenAI. Requires strong infra, distributed systems, and Python/C++/Rust experience.
230k – 385k/yr
Hybrid5+ YOEDevOps / SRE
Operations Engineer, BizTech
AirbnbUnited States
Operations Engineer using AI, LLMs, and intelligent automation to triage tickets, accelerate incident response, build self-healing observability, and automate repetitive operational work in Airbnb's BizTech Global Operations team.
136k – 160k/yr
Remote3+ YOEDevOps / SRE
Systems Integration Engineer, Build Systems | Consumer Devices
OpenAISan Francisco, CA
Build and evolve Bazel, Yocto, and Buildkite-based CI systems for OpenAI consumer device software. Focus on hermetic builds, remote caching, test optimization, observability, and AI-powered failure analysis to accelerate reliable shipping. Requires 5+ years building developer infrastructure at scale.
293k – 325k/yr
Hybrid5+ YOEDevOps / SRE
Software Engineer, CI Platform Infrastructure
AirbnbUnited States
Build and optimize a next-generation CI platform infrastructure for workflow orchestration, scheduling, caching, and autoscaling to accelerate software development for engineers and AI coding agents at scale. Requires interest in distributed systems and knowledge of Kubernetes, EC2, Golang, and Docker.