Software Engineer, GPU Infrastructure
Build and operate scalable GPU/TPU HPC infrastructure for training and serving frontier AI models. The role partners with AI researchers, optimizes distributed workloads across clouds, and requires expertise in Kubernetes, Python, Go, Linux, and high-performance networking.
About the job
Responsibilities
- Build and scale ML-optimized HPC infrastructure, including Kubernetes-based GPU/TPU superclusters across multiple clouds.
- Optimize AI/ML training infrastructure for cost efficiency, reliability, and performance using RDMA, NCCL, and high-speed interconnects.
- Troubleshoot infrastructure bottlenecks, performance degradation, and system failures.
- Design self-service interfaces and workflows for researchers to monitor, debug, and optimize training jobs.
- Translate emerging needs involving JAX, PyTorch, and distributed training into scalable infrastructure solutions.
- Promote observability, automation, and infrastructure-as-code practices.
- Mentor engineers through code reviews, documentation, and cross-team collaboration.
- Participate in a compensated 24x7 on-call rotation.
Requirements
- Deep experience with ML/HPC infrastructure, GPU/TPU clusters, distributed training frameworks, and high-performance computing environments.
- Experience deploying, managing, and troubleshooting Kubernetes clusters for AI workloads.
- Proficiency in Python and Go.
- Familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads.
- Experience collaborating with AI researchers or ML engineers on infrastructure challenges.
- Ability to independently identify bottlenecks, propose solutions, and drive impact in a fast-paced environment.
Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Comprehensive health and dental benefits, including a separate mental-health budget.
- RRSP matching, 401K, or pension scheme.
- Up to six months of 100% parental-leave top-up for either parent.
- Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation.
- Travel budget for remote employees visiting other offices and an annual company offsite.
- Coworking benefit and a $500 home-office stipend.
Skills
Kubernetes, Gpu Clusters, Tpu Clusters, Python, Go, JAX, PyTorch, TensorFlow, Rdma, Nccl, Linux, Distributed Training, High-Performance Computing, Infrastructure As Code, Observability
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.
Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.