Software Engineer, Fleet Infrastructure
Designs, implements, and operates infrastructure systems for model training and deployment on a massive GPU fleet. Requires experience with hyperscale compute, Kubernetes, public clouds like Azure, and strong programming skills.
About the job
Responsibilities
- Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems.
- Interface with researchers and product teams to understand workload requirements.
- Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service.
Requirements
- Experience with hyperscale compute systems.
- Strong programming skills.
- Experience working in public clouds (especially Azure).
- Experience working in Kubernetes.
- Execution focused mentality paired with a rigorous focus on user requirements.
Nice-to-haves
- Understanding of AI/ML workloads.
Skills
Kubernetes, Azure, CI/CD, Job Scheduling, Cluster Management, Snapshot Delivery, Hyperscale Compute, Ai/Ml Workloads
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.