Skip to content
SpellbrushSpellbrush

HPC/ML Infrastructure Engineer

Lead bringup, administration, and operations for a large-scale anime AI training GPU cluster. Bridge researchers and bare-metal systems using SLURM, filesystems, networking, and Linux sysadmin skills.

About the job

Responsibilities

  • Lead bringup, administration, and operations on a large-scale anime AI training cluster
  • Ensure SLURM jobs are running, parallel filesystems are serving, network is transmitting, and models are training
  • Bridge between researchers and bare GPU machines

Requirements

  • Experience with modern HPC software landscape including SLURM (Slinky on K8s), provisioning (warewulf/MAAS/ansible), filesystems (WEKA/VAST/Ceph), VPN/access (tailscale), monitoring (Grafana/Prometheus)
  • Traditional Linux sysadmin skills including LDAP, dmesg triage, and directory permissions
  • Comfortable with physical hardware: racking, stacking, provisioning HGX-based nodes, VLAN design, fiber routing
  • Willing to work on small, fast-paced teams directly with AI researchers
  • On-site collaboration preferred in Tokyo (Akihabara) or San Francisco (Dogpatch); Bay Area strongly preferred due to physical hardware presence
  • Visa sponsorship available

Nice-to-Haves

  • Love of anime and anime aesthetic

Skills

Slurm, Kubernetes, Ansible, Weka, Ceph, Grafana, Prometheus, Linux, Ldap, Hpc

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.