Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.
240k – 300k/yr
Remote8+ YOEDevOps / SRE
About the role
Responsibilities
Build, operate, and troubleshoot production Kubernetes/EKS clusters.
Perform Kubernetes upgrades, node rollouts, and cluster maintenance.
Build and manage AWS infrastructure, including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage.
Define and maintain infrastructure using Terraform.
Build and operate CI/CD and deployment infrastructure.
Troubleshoot production issues across Kubernetes, AWS, Linux, networking, and databases.
Build monitoring, alerting, and observability for critical infrastructure.
Participate in on-call rotations and respond to production incidents.
Identify and solve infrastructure scaling and reliability problems.
Automate operational work using Python, Go, or similar languages.
Help expand infrastructure across new regions and deployment environments.
Requirements
8+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps, Production Engineer, or equivalent infrastructure professional.
Strong hands-on experience operating Kubernetes, including production cluster management rather than only deploying applications to existing clusters.
Experience managing Kubernetes/EKS upgrades and production clusters.
Strong AWS fundamentals, including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases.
Production experience with Terraform or similar infrastructure-as-code tooling.
Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins.
Experience diagnosing production infrastructure and networking problems.
Experience solving meaningful scaling or reliability challenges.
U.S. person status and the ability to access controlled or restricted information as required for the position.
Nice-to-haves
Helm and GitOps experience.
Datadog or similar observability tooling.
PostgreSQL/database operations experience.
Multi-region infrastructure experience.
On-premises or disconnected deployment experience.
Streaming or high-throughput distributed systems experience.
Compensation and Benefits
Annual base salary: $240,000-$300,000.
Equity in the form of stock options.
Comprehensive benefits package, including group health insurance plans.
Paid vacation time, sick leave, holiday pay, and a 401(k) savings plan.
Relocation assistance may be available for eligible roles.
Build and own software state machines and control planes that automate the full lifecycle of GPU infrastructure from bare metal provisioning to running AI inference clusters. Requires strong software engineering experience with orchestration, reconciliation loops, and event-driven systems.
240k – 280k/yrOn-site7+ YOEDevOps / SRE
Staff Platform Engineer, Service Infrastructure
Together AISan Francisco, CA
Staff Platform Engineer owning service infrastructure strategy for Together AI's Product Foundations (API, UI, Billing, IAM). Lead Kubernetes, AWS, Terraform, networking, and reusable primitives to improve reliability, consistency, and scalability across teams.
240k – 280k/yrOn-site7+ YOEDevOps / SRE
Staff Software Engineer, AI Developer Tooling
SentrySan Francisco, CA
Build and own AI-assisted coding domain within Platform Engineering. Audit and expose internal systems via APIs for AI agents, create harness tooling and feedback loops for high-quality AI-generated PRs, automate routine engineering work, and design internal productivity tools. Requires 10+ years software engineering experience and hands-on work with AI coding tools.
240k – 320k/yrHybrid10+ YOEDevOps / SRE
Staff Software Engineer, Site Reliability Engineer (SRE)
HarveySan Francisco, CA
Staff SRE ensures reliability, scalability, and performance of legal AI platform across global regions. Leads incident management, automates operations, optimizes infrastructure costs, and mentors teams. Requires 10+ years SRE experience, IaC, cloud platforms, and observability tools.
Leads architecture and development of scalable managed Kubernetes and AI orchestration systems, providing technical direction for cloud infrastructure reliability and performance. Requires 10+ years in software engineering with deep expertise in Go, Kubernetes, and large-scale systems.