Production Engineer, Compute
Own end-to-end health, repair automation, and qualification of a hyperscale GPU/TPU compute fleet. Build metrics pipelines, firmware tooling, and self-healing repair workflows across Kubernetes and bare metal.
About the job
Role Scope
- Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU and TPU in production — across Kubernetes-orchestrated workloads and bare metal, at scale.
- Turn repair into a pipeline, not a procedure. Build and own the automation that takes a compute failure from detection through triage, parts management, and return to service.
- Design and expand the XPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU and TPU generation.
- Own Redfish and BMC tooling. Firmware-level telemetry, log collection at fleet scale, and the low-level access layer that repair automation and health tooling depend on.
- Own end-to-end reliability, scalability, and operation of the compute fleet at-scale.
What We're Looking For
- Treat toil as a bug. Manual steps in a repair workflow are a backlog item, not a job description.
- Instinct for hardware. Comfortable reasoning about failure modes at the firmware and silicon level.
- Move toward ambiguity, not away from it. Walk into the fog, build the map, and explain it to everyone else.
- Learn at a steep slope. Reach real competence in an unfamiliar domain fast.
- Carry a pager without flinching. Run the incident, write the postmortem, fix the systemic cause, and move on.
- Fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks; drive Claude Code, Cursor, or similar every day.
- Shipped production automation that other teams depend on, and comfortable in any language using AI coding tools.
Bonus
- Hardware lifecycle management and RMA automation.
- BMC/Redfish or IPMI tooling.
- GPU/TPU qualification or burn-in frameworks.
- Workflow and orchestration engines (Temporal, Cadence).
- Metrics and alerting pipelines (Prometheus, Grafana).
- Go or Python.
Salary & Benefits
- Competitive total compensation package (salary + equity).
- Retirement or pension plan, in line with local norms.
- Health, dental, and vision insurance.
- Generous PTO policy, in line with local norms.
- Base salary range: $175,000 - $300,000 per year, depending on experience, skills, qualifications, and location. Total compensation may also include equity in the form of stock options.
Skills
Kubernetes, Redfish, Bmc, Prometheus, Grafana, Go, Python, Temporal, Cadence, Ipmi
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.