Production Engineer owning observability, control plane APIs, fleet state, and new hardware integration for hyperscale GPU infrastructure at an AI compute company. Requires strong automation mindset, on-call experience, and fluency with AI coding tools; distributed systems and observability stack experience preferred.
208k – 269k/yr
On-site5+ YOEDevOps / SRE
About the role
Responsibilities
Own the observability platform: build and operate data pipelines, decoration and correlation engine, and healthcheck framework that make the fleet legible from site down to device and link.
Define and build the API surface for infrastructure: design contracts between production infrastructure and every tool that touches it.
Build the production control plane: unified machine management, actual state inspection, distributed command execution, and the Kubernetes-based infrastructure that underpins it.
Own fleet state as source of truth: SLOs, site lifecycle state, and integration with internal infrastructure management and customer-facing operations platforms.
Land new hardware into the platform cleanly: ZTP, DHCP, DNS, artifacts for every new XPU generation and site integration.
Requirements
Treat toil as a bug and automate repetitive tasks.
Design APIs that age well and avoid leaky abstractions at scale.
Move toward ambiguity, build the map, and explain it to others.
Learn at a steep slope and reach competence in unfamiliar domains quickly.
Comfortably carry a pager, run incidents, write postmortems, and fix systemic causes.
Fluent with AI tooling including LLM APIs, MCP servers, agentic frameworks; drive Claude Code, Cursor, or similar daily.
Shipped production services that other teams depend on at scale; comfortable in any language using AI coding tools.
Nice-to-Haves
Distributed systems and data pipeline engineering.
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.
208k – 269k/yr
On-site5+ YOEDevOps / SRE
Site Reliability Engineer, Compute
FluidstackSan Francisco, CA +3
Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.
208k – 269k/yr
On-site5+ YOEDevOps / SRE
Distributed Systems Engineer
FluidstackSan Francisco, CA +3
Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.
208k – 269k/yr
On-site5+ YOEDevOps / SRE
Software Engineer, Product Infrastructure
NotionSan Francisco, CA +1
Build core frameworks and abstractions for data handling, developer productivity, and performance across Notion's stack. Solve complex challenges like content graph traversal and permission scaling using tools like AWS, Postgres, Node.js, TypeScript, and React.
209k – 240k/yr
HybridDevOps / SRE
Software Engineer, Delivery / CD
OpenAISan Francisco, CA +1
Builds and operates continuous deployment platforms for safe, rapid code rollouts across Kubernetes clusters and global regions. Focuses on progressive delivery, GitOps, automation, and AI-assisted workflows to boost developer productivity.