# Senior ML Systems Engineer, Frameworks & Tooling

**Company:** [Cohere](https://hotfix.jobs/companies/cohere)
**Location:** Remote
**Role:** ML Engineering
**Experience:** 5+ years
**Skills:** JAX, PyTorch, Deepspeed, Megatron, Xformers, CUDA, Nccl, Fsdp, Zero, Kubernetes, Ray, Slurm, Docker, vLLM
**Posted:** 2025-12-01

> Build and evolve the distributed training framework and tooling powering frontier-scale language models. The role focuses on large-scale ML systems, HPC infrastructure, performance optimization, reliability, and developer tooling across multi-node GPU clusters.

## Job Description

## Responsibilities
- Build and own the training framework responsible for large-scale LLM training.
- Design distributed training abstractions, including data, tensor, and pipeline parallelism, FSDP/ZeRO strategies, memory management, and checkpointing.
- Improve training throughput and stability on multi-node clusters, including GB200/300, AMD, and H200/100 systems.
- Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
- Collaborate with infrastructure teams to ensure cluster, container, and hardware configurations support high-performance training.
- Investigate and resolve performance bottlenecks across the ML systems stack.
- Build robust systems for reproducible, debuggable, large-scale runs.
- Build high-performance data loading and caching pipelines.
- Implement performance profiling across the ML systems stack.
- Develop internal metrics and monitoring for training runs.
- Build reproducibility and regression-testing infrastructure.
- Develop performant, fault-tolerant distributed checkpointing systems.

## Requirements
- Strong engineering experience in large-scale distributed training or HPC systems.
- Deep familiarity with JAX internals, distributed training libraries, or custom kernels and fused operations.
- Experience with multi-node cluster orchestration.
- Comfort debugging performance issues across CUDA/NCCL, networking, I/O, and data pipelines.
- Experience with containerized environments.
- Track record of building tools that increase developer velocity for ML teams.
- Excellent judgment regarding performance versus complexity and research velocity versus maintainability.
- Strong collaboration skills across infrastructure, research, and deployment teams.

## Nice to Have
- Experience training LLMs or other large transformer architectures.
- Contributions to ML frameworks such as PyTorch, JAX, DeepSpeed, Megatron, or xFormers.
- Familiarity with evaluation and serving frameworks such as vLLM and TensorRT-LLM.
- Experience with data pipeline optimization, sharded datasets, or caching strategies.
- Background in performance engineering, profiling, or low-level systems.
- Papers at top-tier venues such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP.

## Compensation & Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Full health and dental benefits, including a separate mental-health budget.
- RRSP matching, 401K, or pension scheme.
- 100% parental-leave top-up for up to six months for either parent.
- Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation.
- Travel budget for remote employees visiting other offices and an annual company offsite.
- Co-working benefit and a $500 home-office stipend for remote employees.

## Similar jobs

- [Senior Machine Learning Operations Engineer](https://hotfix.jobs/jobs/102dd624-6c60-41dd-823e-299c080f8b33) - Mercury - Remote - $157k – $208k/yr
- [Senior Software Engineer](https://hotfix.jobs/jobs/c1cce4bb-f642-43f8-9436-a7d1f7f3fbfe) - Traba - New York, NY - $200k – $240k/yr
- [Senior Applied AI Engineer](https://hotfix.jobs/jobs/4a90a405-edd3-40e0-9afb-1c5b87cfd867) - Front - San Francisco, CA - $205k – $238k/yr
- [Senior AI Engineer, Agentic Data Enrichment](https://hotfix.jobs/jobs/5246cc35-17f3-4041-a991-5cda61fcb0cf) - Baselayer - San Francisco, CA - $230k – $340k/yr
- [Senior AI Engineer](https://hotfix.jobs/jobs/4614c8ea-6d88-4da9-bf76-4d241c80e799) - Blee - San Francisco, CA - $150k – $275k/yr

**Apply:** https://hotfix.jobs/jobs/d51f1c6e-0c4a-4e22-977e-8a0a3b960b1c
**Canonical:** https://hotfix.jobs/jobs/d51f1c6e-0c4a-4e22-977e-8a0a3b960b1c