Software Engineer, Compute
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.
About the job
Responsibilities
- Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production — across Kubernetes-orchestrated workloads and bare metal, at scale.
- Turn deployment/repair into a pipeline, not a procedure. Build and own the automation that takes a compute failure from detection through triage, parts management, and return to service. No one-off scripts, no heroics.
- Design and expand the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU generation. Define what "good" looks like before hardware goes into production.
- Own Redfish and BMC tooling. Firmware-level telemetry, log collection at fleet scale, and the low-level access layer that repair automation and health tooling depend on.
- Own end-to-end reliability, scalability, and operation of the compute fleet at-scale. Build aggressive automation, tooling, and incident discipline for one of the largest GPU fleets in the world.
Requirements
- Treat toil as a bug. Manual steps in a repair workflow are a backlog item, not a job description.
- Instinct for hardware. Comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it.
- Move toward ambiguity, not away from it. Walk into the fog, build the map, and explain it to everyone else.
- Learn at a steep slope. Reach real competence in an unfamiliar domain fast.
- Carry a pager without flinching. Run the incident, write the postmortem, fix the systemic cause, and move on.
- Fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks; drive Claude Code, Cursor, or similar every day.
- Shipped production automation that other teams depend on, and comfortable in any language using AI coding tools.
Nice-to-Haves
- Hardware lifecycle management and RMA automation.
- BMC/Redfish or IPMI tooling.
- GPU qualification or burn-in frameworks.
- Workflow and orchestration engines (Temporal, Cadence).
- Metrics and alerting pipelines (Prometheus, Grafana).
- Go or Python.
Skills
Kubernetes, Redfish, Bmc, Prometheus, Grafana, Go, Python, Temporal, Cadence, LLM APIs, Gpu Qualification
Similar jobs
DevOps / SRE jobsBuild and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.