Skip to content
FluidstackFluidstack

Backend Engineer

Production Engineer owning observability, control plane APIs, fleet state, and new hardware integration for hyperscale GPU infrastructure at an AI compute company. Requires strong automation mindset, on-call experience, and fluency with AI coding tools; distributed systems and observability stack experience preferred.

About the job

Responsibilities

  • Own the observability platform: build and operate data pipelines, decoration and correlation engine, and healthcheck framework that make the fleet legible from site down to device and link.
  • Define and build the API surface for infrastructure: design contracts between production infrastructure and every tool that touches it.
  • Build the production control plane: unified machine management, actual state inspection, distributed command execution, and the Kubernetes-based infrastructure that underpins it.
  • Own fleet state as source of truth: SLOs, site lifecycle state, and integration with internal infrastructure management and customer-facing operations platforms.
  • Land new hardware into the platform cleanly: ZTP, DHCP, DNS, artifacts for every new XPU generation and site integration.

Requirements

  • Treat toil as a bug and automate repetitive tasks.
  • Design APIs that age well and avoid leaky abstractions at scale.
  • Move toward ambiguity, build the map, and explain it to others.
  • Learn at a steep slope and reach competence in unfamiliar domains quickly.
  • Comfortably carry a pager, run incidents, write postmortems, and fix systemic causes.
  • Fluent with AI tooling including LLM APIs, MCP servers, agentic frameworks; drive Claude Code, Cursor, or similar daily.
  • Shipped production services that other teams depend on at scale; comfortable in any language using AI coding tools.

Nice-to-Haves

  • Distributed systems and data pipeline engineering.
  • Time-series observability stacks (Prometheus, Thanos, VictoriaMetrics).
  • API design and versioning at scale.
  • Workflow and orchestration engines (Temporal, Cadence).
  • BMC/Redfish or hardware telemetry.
  • Go, Python, and Postgres.

Skills

Kubernetes, Prometheus, Thanos, Victoriametrics, Temporal, Cadence, Go, Python, Postgres, Redfish, LLM APIs, Ai Coding Tools

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.