Skip to content
FluidstackFluidstack

Software Engineer, Compute

Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.

About the job

Responsibilities

  • Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production — across Kubernetes-orchestrated workloads and bare metal, at scale.
  • Turn deployment/repair into a pipeline, not a procedure. Build and own the automation that takes a compute failure from detection through triage, parts management, and return to service. No one-off scripts, no heroics.
  • Design and expand the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU generation. Define what "good" looks like before hardware goes into production.
  • Own Redfish and BMC tooling. Firmware-level telemetry, log collection at fleet scale, and the low-level access layer that repair automation and health tooling depend on.
  • Own end-to-end reliability, scalability, and operation of the compute fleet at-scale. Build aggressive automation, tooling, and incident discipline for one of the largest GPU fleets in the world.

Requirements

  • Treat toil as a bug. Manual steps in a repair workflow are a backlog item, not a job description.
  • Instinct for hardware. Comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it.
  • Move toward ambiguity, not away from it. Walk into the fog, build the map, and explain it to everyone else.
  • Learn at a steep slope. Reach real competence in an unfamiliar domain fast.
  • Carry a pager without flinching. Run the incident, write the postmortem, fix the systemic cause, and move on.
  • Fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks; drive Claude Code, Cursor, or similar every day.
  • Shipped production automation that other teams depend on, and comfortable in any language using AI coding tools.

Nice-to-Haves

  • Hardware lifecycle management and RMA automation.
  • BMC/Redfish or IPMI tooling.
  • GPU qualification or burn-in frameworks.
  • Workflow and orchestration engines (Temporal, Cadence).
  • Metrics and alerting pipelines (Prometheus, Grafana).
  • Go or Python.

Skills

Kubernetes, Redfish, Bmc, Prometheus, Grafana, Go, Python, Temporal, Cadence, LLM APIs, Gpu Qualification

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.