Skip to content
FluidstackFluidstackSan Francisco, CA

Production Engineering Lead, Compute

Lead the compute production engineering team at Fluidstack, owning availability SLOs, automation for node lifecycle (provisioning to remediation), and on-call models for tens of thousands of GPUs at nation-scale. Requires prior leadership of large-fleet SRE/production teams, proven availability improvements, and automated remediation experience.

269k – 335k/yr
On-site7+ YOEEngineering Management

About the role

Lead the Compute Production Engineering Team

  • Own fleet availability for compute: define SLOs, build tooling, and improve the availability number.
  • Build automation for the node lifecycle: provisioning, health checks, remediation, return to service with zero human touch.
  • Set the on-call and escalation model that maintains sharp response times without burning out the team.
  • Lead the team responsible for keeping tens of thousands of GPUs serving customers at nation-scale.

Requirements

  • Led SRE or production engineering teams running large fleets.
  • Demonstrated ability to move an availability number with clear explanations of methods used.
  • Shipped automated remediation that retired a runbook.
  • Experience hiring and growing strong engineers (with positive feedback from reports).

Nice-to-Haves

  • Experience with GPU or HPC fleets.
  • Kubernetes or Slurm.
  • Hardware failure analytics.
  • Customer-facing reliability work.

Skills

SREproduction engineeringKubernetesslurmGPUhpchardware failure analytics
Fluidstack

Software Engineer Tech Lead

FluidstackAustin, TX +3

Tech lead setting technical direction and building core pieces of an internal platform for AI compute infrastructure, including orchestration, integrations, and shared systems. Requires experience designing platforms others depend on, high-bar code review, and daily use of AI coding tools.

269k – 335k/yr
On-site7+ YOEEngineering Management
Fluidstack

Software Engineer Team Lead

FluidstackAustin, TX +3

Lead a small high-velocity engineering team building internal operational software and tools for AI compute infrastructure buildout. Stay hands-on coding while owning priorities, delivery, engineer growth, hiring, and cross-functional product decisions.

269k – 335k/yr
On-site5+ YOEEngineering Management
Clubhouse

Senior Engineering Manager

ClubhouseUnited States

Lead engineering teams building backend, infrastructure, and AI systems for a social product studio. Own people management, delivery, and technical direction as a player-coach with 7+ years engineering and 2+ years management experience.

267k – 315k/yr
Remote7+ YOEEngineering Management
Confluent

Senior Engineering Manager, Flink Control Plane

ConfluentCalifornia

Lead a team of engineers to develop and execute the roadmap for the Flink Control Plane, focusing on architectural excellence, reliability, and scaling for Confluent's managed Flink offering. This role involves significant technical strategy and people leadership.

272k – 319k/yr
Remote10+ YOEEngineering Management
Harvey

Senior Engineering Manager, Model Infrastructure

HarveySan Francisco, CA

Lead the Model Infrastructure engineering team at Harvey to build reliable, scalable platforms for multi-provider AI model operations, routing, observability, and future training infrastructure. Requires 8+ years software engineering experience including multiple years managing teams on large-scale distributed systems.

272k – 355k/yr
Hybrid8+ YOEEngineering Management