Skip to content
Ai2Ai2

Senior Software Engineer, AI Infrastructure

Senior engineer building and operating large-scale HPC infrastructure for AI model training. Owns job scheduling, automation, and performance optimization across GPU clusters.

About the job

Responsibilities

  • Independently design and deliver critical systems spanning the full stack—from the Beaker job scheduler to the execution runtime
  • Build innovative tooling and software-defined infrastructure to accelerate researcher velocity and automate cluster health management
  • Conduct root-cause analysis on complex distributed system failures and implement optimizations for distributed workloads
  • Provide input into the roadmap for managing large-scale HPC systems, including deployment of compute, networking, and storage
  • Review code/design docs, mentor team members, and drive process improvements
  • Communicate and collaborate with internal research staff to share system designs and support implementation

Requirements

  • 8+ years of professional experience developing business-critical software and operating large-scale compute infrastructure
  • Proficiency in Go and/or Python
  • Bachelor’s degree in related field (advanced degree may substitute for experience)
  • Expert-level knowledge of Linux internals and container runtimes (Docker)
  • Proven track record designing, debugging, and optimizing high-scale distributed systems and databases
  • Exceptional writing skills and ability to drive consensus across researchers and engineers
  • Principled approach to engineering and excitement for non-profit research environment

Nice-to-Haves

  • Experience with workload schedulers (Kubernetes, Slurm) and high-performance networking (NCCL, InfiniBand)
  • Prior experience training or fine-tuning frontier AI models
  • Deep systems administration or SRE background in HPC context
  • Contributions to open-source infrastructure or orchestration projects
  • Familiarity with on-prem storage systems (WEKA, Ceph)

Skills

Go, Python, Linux, Docker, Distributed Systems, Kubernetes, Slurm, Nccl, InfiniBand, SRE

PrizePicks

PrizePicks

United States

Senior Site Reliability Engineer
$120k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for designing and operating reliable, scalable production infrastructure, leading incident response, and improving observability and resilience. Requires 5+ years of reliability-focused engineering experience and expertise across cloud, infrastructure as code, Kubernetes, monitoring, and application development.

Axle

Axle

Rockville, MD

Lead Scientific Imaging Systems Engineer
$120k+/yrOn-site7+ YOEDevOps / SRE

Leads end-to-end infrastructure for a scientific imaging platform, covering Linux administration, GPU/HPC systems, storage, upgrades, and vendor coordination. The role supports AI-enabled imaging workflows and requires extensive production Linux, Image Artist, GPU, HPC, and enterprise storage experience.

MongoDB

MongoDB

Palo Alto, CA
Senior Network Engineer
$118k+/yrHybrid6+ YOEDevOps / SRE

Senior network engineer responsible for designing, operating, and securing MongoDB’s global network and VPN infrastructure. The role requires 6+ years of networking or systems engineering experience, strong enterprise networking expertise, automation skills, and the ability to lead complex infrastructure initiatives.

Upstart

Upstart

United States

Senior DevOps Engineer
$136k+/yrRemote3+ YOEDevOps / SRE

Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.

Bloomerang

Bloomerang

United States

Senior Software Engineer, Site Reliability
$115k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.