# Backend Engineer

**Company:** [Fluidstack](https://hotfix.jobs/companies/fluidstack)
**Location:** San Francisco, CA, New York, NY, Austin, TX, Seattle, WA
**Role:** DevOps / SRE
**Salary:** $208k – $269k/yr
**Experience:** 5+ years
**Skills:** Kubernetes, Prometheus, thanos, victoriametrics, Temporal, cadence, Go, Python, Postgres, redfish, LLM APIs, ai coding tools
**Posted:** 2026-07-05

> Production Engineer owning observability, control plane APIs, fleet state, and new hardware integration for hyperscale GPU infrastructure at an AI compute company. Requires strong automation mindset, on-call experience, and fluency with AI coding tools; distributed systems and observability stack experience preferred.

## Job Description

## Responsibilities
- Own the observability platform: build and operate data pipelines, decoration and correlation engine, and healthcheck framework that make the fleet legible from site down to device and link.
- Define and build the API surface for infrastructure: design contracts between production infrastructure and every tool that touches it.
- Build the production control plane: unified machine management, actual state inspection, distributed command execution, and the Kubernetes-based infrastructure that underpins it.
- Own fleet state as source of truth: SLOs, site lifecycle state, and integration with internal infrastructure management and customer-facing operations platforms.
- Land new hardware into the platform cleanly: ZTP, DHCP, DNS, artifacts for every new XPU generation and site integration.

## Requirements
- Treat toil as a bug and automate repetitive tasks.
- Design APIs that age well and avoid leaky abstractions at scale.
- Move toward ambiguity, build the map, and explain it to others.
- Learn at a steep slope and reach competence in unfamiliar domains quickly.
- Comfortably carry a pager, run incidents, write postmortems, and fix systemic causes.
- Fluent with AI tooling including LLM APIs, MCP servers, agentic frameworks; drive Claude Code, Cursor, or similar daily.
- Shipped production services that other teams depend on at scale; comfortable in any language using AI coding tools.

## Nice-to-Haves
- Distributed systems and data pipeline engineering.
- Time-series observability stacks (Prometheus, Thanos, VictoriaMetrics).
- API design and versioning at scale.
- Workflow and orchestration engines (Temporal, Cadence).
- BMC/Redfish or hardware telemetry.
- Go, Python, and Postgres.

## Similar roles

- [Software Engineer, Compute](https://hotfix.jobs/jobs/a47fd61c-f69f-4b10-ad03-bd207719bb21) - Fluidstack - San Francisco, CA - $208k – $269k/yr
- [Site Reliability Engineer, Compute](https://hotfix.jobs/jobs/4374777f-3589-4aac-bd7c-61c791f5a467) - Fluidstack - San Francisco, CA - $208k – $269k/yr
- [Distributed Systems Engineer](https://hotfix.jobs/jobs/8505fdd9-bfbf-438a-9c6f-64b2d80f7c5c) - Fluidstack - San Francisco, CA - $208k – $269k/yr
- [Software Engineer, Product Infrastructure](https://hotfix.jobs/jobs/d6939038-b51c-4151-a009-37866deb220b) - Notion - San Francisco, CA - $209k – $240k/yr
- [Software Engineer, Delivery / CD](https://hotfix.jobs/jobs/187d0ae4-8435-4cac-a1b6-6259aeb7c077) - OpenAI - San Francisco, CA - $210k – $490k/yr

**Apply:** https://hotfix.jobs/jobs/c39aa8f3-f056-4424-b3ef-ba98f7497ef9
**Canonical:** https://hotfix.jobs/jobs/c39aa8f3-f056-4424-b3ef-ba98f7497ef9