Senior Software Engineer, AI Infrastructure
Senior engineer building and operating large-scale HPC infrastructure for AI model training. Owns job scheduling, automation, and performance optimization across GPU clusters.
About the job
Responsibilities
- Independently design and deliver critical systems spanning the full stack—from the Beaker job scheduler to the execution runtime
- Build innovative tooling and software-defined infrastructure to accelerate researcher velocity and automate cluster health management
- Conduct root-cause analysis on complex distributed system failures and implement optimizations for distributed workloads
- Provide input into the roadmap for managing large-scale HPC systems, including deployment of compute, networking, and storage
- Review code/design docs, mentor team members, and drive process improvements
- Communicate and collaborate with internal research staff to share system designs and support implementation
Requirements
- 8+ years of professional experience developing business-critical software and operating large-scale compute infrastructure
- Proficiency in Go and/or Python
- Bachelor’s degree in related field (advanced degree may substitute for experience)
- Expert-level knowledge of Linux internals and container runtimes (Docker)
- Proven track record designing, debugging, and optimizing high-scale distributed systems and databases
- Exceptional writing skills and ability to drive consensus across researchers and engineers
- Principled approach to engineering and excitement for non-profit research environment
Nice-to-Haves
- Experience with workload schedulers (Kubernetes, Slurm) and high-performance networking (NCCL, InfiniBand)
- Prior experience training or fine-tuning frontier AI models
- Deep systems administration or SRE background in HPC context
- Contributions to open-source infrastructure or orchestration projects
- Familiarity with on-prem storage systems (WEKA, Ceph)
Skills
Go, Python, Linux, Docker, Distributed Systems, Kubernetes, Slurm, Nccl, InfiniBand, SRE
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for designing and operating reliable, scalable production infrastructure, leading incident response, and improving observability and resilience. Requires 5+ years of reliability-focused engineering experience and expertise across cloud, infrastructure as code, Kubernetes, monitoring, and application development.
Leads end-to-end infrastructure for a scientific imaging platform, covering Linux administration, GPU/HPC systems, storage, upgrades, and vendor coordination. The role supports AI-enabled imaging workflows and requires extensive production Linux, Image Artist, GPU, HPC, and enterprise storage experience.
Senior network engineer responsible for designing, operating, and securing MongoDB’s global network and VPN infrastructure. The role requires 6+ years of networking or systems engineering experience, strong enterprise networking expertise, automation skills, and the ability to lead complex infrastructure initiatives.
Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.
Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.