Staff Site Reliability Engineer – Automation and Platform
Leads automation and platform engineering for ultra-reliable AI inference infrastructure, architecting self-service GitOps pipelines, observability, and tooling to eliminate toil across datacenters. Requires 8+ years SRE experience with large-scale clusters and tools like Argo CD and Prometheus.
About the job
Key Responsibilities
- Define and implement a robust strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions.
- Architect self-service platforms and internal tooling that let product teams, external customers, and cluster operators safely trigger and observe critical workflows with minimal handoffs.
- Define and evolve reliability practices for inference workloads, including SLOs and SLIs for latency, throughput, and accuracy stability; error budgets; blameless postmortems; chaos testing; and capacity forecasting across multi-datacenter and on-prem environments.
- Mentor mid-level SREs, support critical incident escalations, and use production pain points to prioritize the highest-leverage automation work.
- Measure and drive impact through clear metrics, including toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
Required Experience & Skills
- 8+ years in SRE, infrastructure engineering, or platform engineering, with a strong record of improving automation and reliability at large scale in FAANG, hyperscaler, or similarly demanding environments.
- Deep expertise operating large scale heterogenous clusters with a proprietary cloud control plane.
- Proven track record designing and delivering CI/CD or GitOps systems using Argo CD or similar tools, with strong safety and observability built in.
- Hands-on experience with observability systems such as Loki, Tempo, Mimir, and Prometheus.
- Ability to lead complex projects end to end, influence cross-functional stakeholders, and communicate technical direction clearly.
Nice-to-Haves
- Experience with Bazel or other large-scale build systems in production.
- Background in AI/ML inference systems, including model serving runtimes, GPU or wafer-scale orchestration, latency and accuracy SLOs, or drift monitoring.
- Prior work on predictive autoscaling, chaos engineering, or cost-aware capacity planning for compute-intensive workloads.
Skills
Argo Cd, GitOps, Prometheus, Loki, Tempo, Mimir, Bazel, CI/CD, SLOs, Chaos Engineering
Similar jobs
DevOps / SRE jobsStaff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.
Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.
Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.