# Software Engineer, GPU Infrastructure

**Company:** [Cohere](https://hotfix.jobs/companies/cohere)
**Location:** Unspecified
**Role:** DevOps / SRE
**Experience:** 7+ years
**Skills:** Kubernetes, Gpu Clusters, Tpu Clusters, Python, Go, JAX, PyTorch, TensorFlow, Rdma, Nccl, Linux, Distributed Training, High-Performance Computing, Infrastructure As Code, Observability
**Posted:** 2026-09-03

> Build and operate scalable GPU/TPU HPC infrastructure for training and serving frontier AI models. The role partners with AI researchers, optimizes distributed workloads across clouds, and requires expertise in Kubernetes, Python, Go, Linux, and high-performance networking.

## Job Description

## Responsibilities
- Build and scale ML-optimized HPC infrastructure, including Kubernetes-based GPU/TPU superclusters across multiple clouds.
- Optimize AI/ML training infrastructure for cost efficiency, reliability, and performance using RDMA, NCCL, and high-speed interconnects.
- Troubleshoot infrastructure bottlenecks, performance degradation, and system failures.
- Design self-service interfaces and workflows for researchers to monitor, debug, and optimize training jobs.
- Translate emerging needs involving JAX, PyTorch, and distributed training into scalable infrastructure solutions.
- Promote observability, automation, and infrastructure-as-code practices.
- Mentor engineers through code reviews, documentation, and cross-team collaboration.
- Participate in a compensated 24x7 on-call rotation.

## Requirements
- Deep experience with ML/HPC infrastructure, GPU/TPU clusters, distributed training frameworks, and high-performance computing environments.
- Experience deploying, managing, and troubleshooting Kubernetes clusters for AI workloads.
- Proficiency in Python and Go.
- Familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads.
- Experience collaborating with AI researchers or ML engineers on infrastructure challenges.
- Ability to independently identify bottlenecks, propose solutions, and drive impact in a fast-paced environment.

## Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Comprehensive health and dental benefits, including a separate mental-health budget.
- RRSP matching, 401K, or pension scheme.
- Up to six months of 100% parental-leave top-up for either parent.
- Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation.
- Travel budget for remote employees visiting other offices and an annual company offsite.
- Coworking benefit and a $500 home-office stipend.

## Similar jobs

- [Senior Platform Engineer](https://hotfix.jobs/jobs/71245322-afd8-4019-8525-65088fed1493) - Shield AI - San Diego, CA - $141k – $212k/yr
- [Senior Network Engineer](https://hotfix.jobs/jobs/d4ecdaa5-c0ab-49f3-baed-0b13deaaf6c0) - Shield AI - San Mateo, CA - $140k – $211k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/3317571d-7a67-4059-b23e-af1a9033cf96) - Astra - Remote - $190k – $230k/yr
- [Senior Software Engineer, Cloud Engineering](https://hotfix.jobs/jobs/e14b1da4-1208-4648-a785-30451d299dd7) - Mozilla - Remote - CA$95k – CA$139k/yr
- [Senior Software Engineer, Cloud Engineering](https://hotfix.jobs/jobs/d332c3cb-7ad5-4a0d-b2c2-af9f684bd506) - Mozilla - Remote

**Apply:** https://hotfix.jobs/jobs/0140c9f5-05ed-4ac4-8f3c-89978a1ea960
**Canonical:** https://hotfix.jobs/jobs/0140c9f5-05ed-4ac4-8f3c-89978a1ea960