Skip to content
FluidstackFluidstackSan Francisco, CA

Distributed Systems Engineer

Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.

208k – 269k/yr
On-site5+ YOEDevOps / SRE

About the role

Role Scope

Own the observability platform. Build and operate the data pipelines, decoration and correlation engine, and healthcheck framework that make the fleet legible — from site down to device and link. No other team should need to scrape production directly to answer a question.

Define and build the API surface for infrastructure. Design the contracts between production infrastructure and every tool that touches it. All other teams at Fluidstack use your tooling to manage and operate our hyperscale fleet.

Build the production control plane. Unified machine management, actual state inspection, distributed command execution — and the Kubernetes-based infrastructure that underpins it all.

Own fleet state as source of truth. SLOs, site lifecycle state, and integration with internal infrastructure management and customer-facing operations platforms. What the system says about itself should match reality, and you're accountable when it doesn't.

Land new hardware into the platform cleanly. ZTP, DHCP, DNS, artifacts — every new XPU generation and site integration goes through IaaS before production.

What We're Looking For

  • You treat toil as a bug. If something requires a human to do it twice, you build the thing that makes it not require a human.
  • You design APIs that age well. You've felt the pain of a leaky abstraction at scale and you don't repeat it.
  • You move toward ambiguity, not away from it. You walk into the fog, build the map, and explain it to everyone else.
  • You learn at a steep slope. You reach real competence in an unfamiliar domain fast. We value this over existing expertise.
  • You carry a pager without flinching. You run the incident, write the postmortem, fix the systemic cause, and move on.
  • You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks, and you drive Claude Code, Cursor, or similar every day.
  • You've shipped production services that other teams depend on at scale, and you're comfortable in any language using AI coding tools.

Bonus:

  • Distributed systems and data pipeline engineering.
  • Time-series observability stacks (Prometheus, Thanos, VictoriaMetrics).
  • API design and versioning at scale.
  • Workflow and orchestration engines (Temporal, Cadence).
  • BMC/Redfish or hardware telemetry.
  • Go, Python, and Postgres.

Skills

KubernetesPrometheusthanosvictoriametricsGoPythonPostgresTemporalcadenceredfishAPI DesignDistributed SystemsData Pipelines

Similar roles

DevOps / SRE jobs
Fluidstack

Software Engineer, Networking

FluidstackSan Francsisco, CA

Build and own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for one of the world's largest AI datacenter networks. Requires systems thinking, on-call ownership, and fluency with AI coding tools plus Go/Python network automation experience.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Software Engineer, Compute

FluidstackSan Francisco, CA +3

Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Site Reliability Engineer, Compute

FluidstackSan Francisco, CA +3

Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Network Software Engineer

FluidstackSan Francisco, CA +3

Build and own end-to-end network monitoring, debugging tooling, repair automation pipelines, and qualification frameworks for Fluidstack's massive AI datacenter fleet. Requires systems thinking, on-call ownership, Go/Python automation experience, network protocols expertise, and daily use of AI coding tools.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Network Automation Engineer

FluidstackSan Francisco, CA +3

Own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for massive AI datacenter infrastructure. Requires systems thinking, building automation in Go/Python, experience with optical networks and protocols like gNMI/gRPC/NETCONF/SONiC, and daily use of AI coding tools.

208k – 269k/yr
On-site5+ YOEDevOps / SRE