Manager, Site Reliability Engineering
Lead and mentor SRE teams managing Okta's IDaaS platform on AWS, driving DevOps maturity, platform reliability, and tooling across Kubernetes, CI/CD, observability, and automation domains.
Build and automate the IaaS layer for large-scale GPU fleets at Fluidstack, including bare-metal provisioning, tenant isolation, lifecycle management, and debugging across compute, network, and storage. Requires production infrastructure automation experience in Go or Python, deep Linux debugging skills, and a software-first approach to operational toil.
Build the IaaS layer: provisioning automation, tenant isolation, and lifecycle management for GPU fleets. Automate bare-metal workflows end to end: image, boot, configure, validate, hand over. Debug across compute, network, and storage when tenant workloads misbehave. Write software that removes toil: every manual runbook is a backlog item with your name on it.
You've built infrastructure automation running in production, in Go or Python against real systems. You've worked with bare-metal provisioning tooling and know where it lies to you. You debug Linux deeply: kernel, drivers, networking. You treat operational pain as a software bug, and you fix it.
Bonus: GPU systems. Kubernetes internals. IPMI and Redfish. Storage systems.
Lead and mentor SRE teams managing Okta's IDaaS platform on AWS, driving DevOps maturity, platform reliability, and tooling across Kubernetes, CI/CD, observability, and automation domains.
Software Engineer building scalable control and data plane infrastructure for Anyscale's Ray platform. Design and optimize cluster orchestration, scheduling, Kubernetes deployments, and accelerator support for distributed AI/ML workloads. Requires 3+ years production experience with cloud-native tech, Go/Python, and distributed systems.
Build and own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for one of the world's largest AI datacenter networks. Requires systems thinking, on-call ownership, and fluency with AI coding tools plus Go/Python network automation experience.
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.
Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.