Skip to content
AnthropicAnthropic

Staff Engineer, Datacenter Server Lifecycle

Owns end-to-end server lifecycle in datacenters at scale, from provisioning to decommissioning, with strong focus on automation, trusted compute security, and hardware operations for AI workloads. Requires hands-on server hardware experience and proficiency in Python/Rust/Go plus cloud infra like Kubernetes/AWS/GCP.

About the job

Key Responsibilities

  • Lead the build-out of automation to support datacenters containing tens of thousands of servers.
  • Define and own the end-to-end server lifecycle strategy — from provisioning and deployment through operation, maintenance, refresh, and decommissioning — and maintain automation and operational procedures for common lifecycle events (e.g., hardware failures, firmware upgrades, fleet rotations).
  • Partner closely with Infrastructure Security to design and enforce trusted compute standards across the server lifecycle.
  • Work closely with our Networking team to ensure end-to-end connectivity across all sites.
  • Build and maintain tooling to track machine health, configuration, and operational status across the full datacenter fleet.

Minimum Qualifications

  • Hands-on experience with server hardware, including rack deployment, cabling, troubleshooting, and understanding failure modes at scale.
  • End-to-end understanding of hardware lifecycle management: asset tracking, provisioning workflows, maintenance scheduling, and decommissioning practices.
  • Proficiency in at least one programming language (e.g., Python, Rust, Go, or Java).
  • Working knowledge of modern cloud infrastructure, including Kubernetes, Infrastructure as Code, AWS, and GCP.
  • Ability to communicate clearly and build consensus with a wide range of stakeholders.
  • Comfort navigating ambiguity and making progress on complex, cross-functional problems.
  • Willingness to travel occasionally to datacenter sites across North America.

Preferred Qualifications

  • 8+ years of experience in datacenter operations, hardware infrastructure management, or a closely related discipline.
  • Hands-on experience with GPU or AI accelerator hardware (e.g., NVIDIA A100/H100, AMD MI300, Google TPUs, or AWS Trainium) and an understanding of their operational demands.
  • Familiarity with modern provisioning tooling such as coreboot, LinuxBoot, or u-root.
  • Experience building or contributing to datacenter automation or fleet management platforms.
  • Experience building and deploying server operating system distributions across thousands of hosts.
  • Background in large-scale capacity planning and hardware refresh strategy, ideally at a hyperscaler or large cloud provider.
  • Experience with trusted compute and hardware security concepts such as secure boot, TPM, hardware attestation, and firmware verification — or a strong desire to develop deep expertise in this area.

Skills

Python, Rust, Go, Java, Kubernetes, Infrastructure As Code, AWS, GCP, Nvidia A100/H100, Amd Mi300, Google Tpus, Aws Trainium, Coreboot, Linuxboot, U-Root

Anthropic

Anthropic

San Francisco, CA
Staff+ Site Reliability Engineer, Safeguards ML Infra
$320k+/yrHybrid8+ YOEDevOps / SRE

Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.

Headway

Headway

San Francisco, CA
Staff Infrastructure Engineer
$265k+/yrRemote8+ YOEDevOps / SRE

Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Deployment Automation Engineer
$250k+/yrOn-site12+ YOEDevOps / SRE

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Software Engineer, DC Infrastructure
$250k+/yrOn-site7+ YOEDevOps / SRE

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.