# Distributed Systems Engineer

**Company:** [Fluidstack](https://hotfix.jobs/companies/fluidstack)
**Location:** San Francisco, CA, New York, NY, Austin, TX, Seattle, WA
**Role:** DevOps / SRE
**Salary:** $208k – $269k/yr
**Experience:** 5+ years
**Skills:** Kubernetes, Prometheus, thanos, victoriametrics, Go, Python, Postgres, Temporal, cadence, redfish, API Design, Distributed Systems, Data Pipelines
**Posted:** 2026-07-05

> Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.

## Job Description

## Role Scope
Own the observability platform. Build and operate the data pipelines, decoration and correlation engine, and healthcheck framework that make the fleet legible — from site down to device and link. No other team should need to scrape production directly to answer a question.

Define and build the API surface for infrastructure. Design the contracts between production infrastructure and every tool that touches it. All other teams at Fluidstack use your tooling to manage and operate our hyperscale fleet.

Build the production control plane. Unified machine management, actual state inspection, distributed command execution — and the Kubernetes-based infrastructure that underpins it all.

Own fleet state as source of truth. SLOs, site lifecycle state, and integration with internal infrastructure management and customer-facing operations platforms. What the system says about itself should match reality, and you're accountable when it doesn't.

Land new hardware into the platform cleanly. ZTP, DHCP, DNS, artifacts — every new XPU generation and site integration goes through IaaS before production.

## What We're Looking For
- You treat toil as a bug. If something requires a human to do it twice, you build the thing that makes it not require a human.
- You design APIs that age well. You've felt the pain of a leaky abstraction at scale and you don't repeat it.
- You move toward ambiguity, not away from it. You walk into the fog, build the map, and explain it to everyone else.
- You learn at a steep slope. You reach real competence in an unfamiliar domain fast. We value this over existing expertise.
- You carry a pager without flinching. You run the incident, write the postmortem, fix the systemic cause, and move on.
- You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks, and you drive Claude Code, Cursor, or similar every day.
- You've shipped production services that other teams depend on at scale, and you're comfortable in any language using AI coding tools.

**Bonus:**
- Distributed systems and data pipeline engineering.
- Time-series observability stacks (Prometheus, Thanos, VictoriaMetrics).
- API design and versioning at scale.
- Workflow and orchestration engines (Temporal, Cadence).
- BMC/Redfish or hardware telemetry.
- Go, Python, and Postgres.

## Similar roles

- [Software Engineer, Networking](https://hotfix.jobs/jobs/c083d3e7-a957-45f7-9406-c00edae189b3) - Fluidstack - San Francsisco, CA - $208k – $269k/yr
- [Software Engineer, Compute](https://hotfix.jobs/jobs/a47fd61c-f69f-4b10-ad03-bd207719bb21) - Fluidstack - San Francisco, CA - $208k – $269k/yr
- [Site Reliability Engineer, Compute](https://hotfix.jobs/jobs/4374777f-3589-4aac-bd7c-61c791f5a467) - Fluidstack - San Francisco, CA - $208k – $269k/yr
- [Network Software Engineer](https://hotfix.jobs/jobs/3691ff92-fa54-43af-8c26-577a95debbae) - Fluidstack - San Francisco, CA - $208k – $269k/yr
- [Network Automation Engineer](https://hotfix.jobs/jobs/98f17fdc-79fd-4af9-a06f-5d58edb06019) - Fluidstack - San Francisco, CA - $208k – $269k/yr

**Apply:** https://hotfix.jobs/jobs/8505fdd9-bfbf-438a-9c6f-64b2d80f7c5c
**Canonical:** https://hotfix.jobs/jobs/8505fdd9-bfbf-438a-9c6f-64b2d80f7c5c