Production Engineering Lead, Compute
Lead the compute production engineering team at Fluidstack, owning availability SLOs, automation for node lifecycle (provisioning to remediation), and on-call models for tens of thousands of GPUs at nation-scale. Requires prior leadership of large-fleet SRE/production teams, proven availability improvements, and automated remediation experience.
About the job
Lead the Compute Production Engineering Team
- Own fleet availability for compute: define SLOs, build tooling, and improve the availability number.
- Build automation for the node lifecycle: provisioning, health checks, remediation, return to service with zero human touch.
- Set the on-call and escalation model that maintains sharp response times without burning out the team.
- Lead the team responsible for keeping tens of thousands of GPUs serving customers at nation-scale.
Requirements
- Led SRE or production engineering teams running large fleets.
- Demonstrated ability to move an availability number with clear explanations of methods used.
- Shipped automated remediation that retired a runbook.
- Experience hiring and growing strong engineers (with positive feedback from reports).
Nice-to-Haves
- Experience with GPU or HPC fleets.
- Kubernetes or Slurm.
- Hardware failure analytics.
- Customer-facing reliability work.
Skills
SRE, Production Engineering, Kubernetes, Slurm, GPU, Hpc, Hardware Failure Analytics
Similar jobs
Engineering Management jobsLeads the mortgage engineering organization, owning platform architecture, delivery, business-line outcomes, and team development. Requires senior engineering management experience, extensive software engineering experience, large-team leadership, business ownership, and expertise in scalable systems and AI.
Leads the Machine Learning Engineering team, setting technical strategy, developing engineers, and shipping reliable ML and agentic AI systems for product intelligence, fraud detection, and workforce integrity. Requires 10+ years building production software and ML/AI systems plus substantial engineering management experience.
Leads and grows a team building shared test systems and tooling for production-representative, workload, performance, and failure-mode validation. Requires engineering management experience, distributed-systems expertise, and a record of delivering widely adopted internal platforms.
Leads machine learning engineering teams responsible for relevance, personalization, and video discovery systems serving Reddit’s core product. Requires 5+ years managing ML teams, hands-on experience with production ML systems, and strong recommender-systems expertise.
Leads a team building retrieval and machine-learned ranking systems that match patients with therapists across a healthcare marketplace. Requires substantial engineering management experience, production software or ML expertise, and hands-on ownership of search, ranking, recommendation, or personalization systems.