Staff Engineer, Datacenter Server Lifecycle
Owns end-to-end server lifecycle in datacenters at scale, from provisioning to decommissioning, with strong focus on automation, trusted compute security, and hardware operations for AI workloads. Requires hands-on server hardware experience and proficiency in Python/Rust/Go plus cloud infra like Kubernetes/AWS/GCP.
About the job
Key Responsibilities
- Lead the build-out of automation to support datacenters containing tens of thousands of servers.
- Define and own the end-to-end server lifecycle strategy — from provisioning and deployment through operation, maintenance, refresh, and decommissioning — and maintain automation and operational procedures for common lifecycle events (e.g., hardware failures, firmware upgrades, fleet rotations).
- Partner closely with Infrastructure Security to design and enforce trusted compute standards across the server lifecycle.
- Work closely with our Networking team to ensure end-to-end connectivity across all sites.
- Build and maintain tooling to track machine health, configuration, and operational status across the full datacenter fleet.
Minimum Qualifications
- Hands-on experience with server hardware, including rack deployment, cabling, troubleshooting, and understanding failure modes at scale.
- End-to-end understanding of hardware lifecycle management: asset tracking, provisioning workflows, maintenance scheduling, and decommissioning practices.
- Proficiency in at least one programming language (e.g., Python, Rust, Go, or Java).
- Working knowledge of modern cloud infrastructure, including Kubernetes, Infrastructure as Code, AWS, and GCP.
- Ability to communicate clearly and build consensus with a wide range of stakeholders.
- Comfort navigating ambiguity and making progress on complex, cross-functional problems.
- Willingness to travel occasionally to datacenter sites across North America.
Preferred Qualifications
- 8+ years of experience in datacenter operations, hardware infrastructure management, or a closely related discipline.
- Hands-on experience with GPU or AI accelerator hardware (e.g., NVIDIA A100/H100, AMD MI300, Google TPUs, or AWS Trainium) and an understanding of their operational demands.
- Familiarity with modern provisioning tooling such as coreboot, LinuxBoot, or u-root.
- Experience building or contributing to datacenter automation or fleet management platforms.
- Experience building and deploying server operating system distributions across thousands of hosts.
- Background in large-scale capacity planning and hardware refresh strategy, ideally at a hyperscaler or large cloud provider.
- Experience with trusted compute and hardware security concepts such as secure boot, TPM, hardware attestation, and firmware verification — or a strong desire to develop deep expertise in this area.
Skills
Python, Rust, Go, Java, Kubernetes, Infrastructure As Code, AWS, GCP, Nvidia A100/H100, Amd Mi300, Google Tpus, Aws Trainium, Coreboot, Linuxboot, U-Root
Similar jobs
DevOps / SRE jobsStaff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.