Infrastructure Engineer (GPU & Compute)

180k – 200kNew York, NYSan Francisco, CASeattle, WADevOps / SRERemote5+ YOEMay 1

Summary

Owns GPU diagnostics, validation workflows, and automation for bare-metal infrastructure supporting AI/ML workloads. Requires 5+ years in systems engineering with strong Linux, Python, and NVIDIA tools expertise.

About the role

What You'll Do

Systems, Image & Validation Infrastructure

Own and evolve systems for image management, deployment, and validation across bare-metal infrastructure
Run and maintain test clusters used for system validation, diagnostics, and bring-up
Validate firmware, drivers, and OS images across compute and GPU-enabled systems
Support hardware qualification efforts for next-generation platforms

GPU Diagnostics & Performance

Own GPU diagnostics and validation workflows across large-scale infrastructure
Diagnose and resolve complex issues across GPUs, drivers, OS, and hardware layers
Analyze system and GPU performance using tools such as NVIDIA DCGM
Identify failure patterns and drive improvements in system stability and validation coverage

Automation & Tooling

Build and maintain automation for provisioning, validation, and system bring-up
Develop Python-based tools and workflows to improve efficiency and reduce manual operational overhead
Improve the reliability, repeatability, and scalability of image pipelines and validation systems

Systems & Operations

Manage and operate Linux-based systems in production and validation environments
Manage virtualization technology
Support bare-metal provisioning workflows, including PXE and image-based systems
Interface with hardware management systems (e.g., IPMI, Redfish) for monitoring and debugging

Cross-Functional Collaboration

Partner with Infrastructure, Hardware, and Data Center teams on system bring-up and validation
Collaborate with platform and ML teams to ensure systems meet workload requirements
Contribute to best practices for provisioning, diagnostics, and lifecycle management of infrastructure

What You'll Need

Required Qualifications

5+ years of experience in infrastructure engineering, systems engineering, or related roles
Strong Linux systems experience in production environments
Hands-on experience with GPU-enabled systems and tools such as NVIDIA DCGM
Familiarity with bare-metal provisioning and system bring-up workflows
Proficiency in Python or similar scripting/programming languages for automation
Ability to debug complex issues across hardware, OS, GPUs, and system software

Ideal Experience

Experience with high-performance interconnects (e.g., InfiniBand, NVLink)
Experience with PXE boot environments, LiveCD systems, or image-based provisioning workflows
Experience with hardware management interfaces such as iDRAC, IPMI, or Redfish
Data center operations experience, including working with physical hardware
Experience supporting AI/ML or HPC workloads at scale
Experience with GPU validation frameworks or large-scale hardware qualification processes

Skills

LinuxPythonNVIDIA DCGMbare-metal provisioningPXEIPMIRedfishInfiniBandNVLinkvirtualization

Similar roles at this salary range

All DevOps / SRE jobs →

Plaid

Jun 19

Staff Site Reliability Engineer, Release Engineering

Staff SRE on the Release Engineering team defining and scaling reliability practices, architecting SLO/error-budget programs, and driving progressive delivery and automated safety gates across product engineering.

208k – 274kNew York, NYDevOps / SREHybrid8+ YOEGoSLO

Fivetran

Jun 18

Senior Site Reliability Engineer

Senior SRE responsible for production infrastructure reliability, incident response, deployment automation, and scaling SaaS systems on Kubernetes and major cloud platforms.

175k – 210kOakland, CADevOps / SREHybrid5+ YOEAWSGCP

Dropbox

Jun 18

Senior Infrastructure Software Engineer, Storage Core

Senior engineer building and operating Dropbox's exabyte-scale distributed storage systems. Focus on replication, erasure coding, performance, and reliability in Go/Rust.

180k – 274kUnited StatesDevOps / SRERemote9+ YOEGoC++

Okta

Jun 17

Staff Site Reliability Engineer - Observability

Staff SRE focused on building and scaling a comprehensive observability platform on GCP using Terraform, Splunk, and Grafana. Requires 5+ years GCP observability experience and strong coding skills in Python or Go.

194k – 267kBellevue, WA +4DevOps / SREHybrid5+ YOEGoGKE

Cribl

Jun 17

Sr Software Engineer, Storage

Senior Software Engineer on the Storage team building autoscaling, self-healing infrastructure-as-code systems that manage petabyte-scale telemetry storage on AWS.

175k – 205kUnited StatesDevOps / SRERemote5+ YOEGoS3

Apply