Skip to content
CohereCohere

Senior ML Systems Engineer, Frameworks & Tooling

Build and evolve the distributed training framework and tooling powering frontier-scale language models. The role focuses on large-scale ML systems, HPC infrastructure, performance optimization, reliability, and developer tooling across multi-node GPU clusters.

About the job

Responsibilities

  • Build and own the training framework responsible for large-scale LLM training.
  • Design distributed training abstractions, including data, tensor, and pipeline parallelism, FSDP/ZeRO strategies, memory management, and checkpointing.
  • Improve training throughput and stability on multi-node clusters, including GB200/300, AMD, and H200/100 systems.
  • Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
  • Collaborate with infrastructure teams to ensure cluster, container, and hardware configurations support high-performance training.
  • Investigate and resolve performance bottlenecks across the ML systems stack.
  • Build robust systems for reproducible, debuggable, large-scale runs.
  • Build high-performance data loading and caching pipelines.
  • Implement performance profiling across the ML systems stack.
  • Develop internal metrics and monitoring for training runs.
  • Build reproducibility and regression-testing infrastructure.
  • Develop performant, fault-tolerant distributed checkpointing systems.

Requirements

  • Strong engineering experience in large-scale distributed training or HPC systems.
  • Deep familiarity with JAX internals, distributed training libraries, or custom kernels and fused operations.
  • Experience with multi-node cluster orchestration.
  • Comfort debugging performance issues across CUDA/NCCL, networking, I/O, and data pipelines.
  • Experience with containerized environments.
  • Track record of building tools that increase developer velocity for ML teams.
  • Excellent judgment regarding performance versus complexity and research velocity versus maintainability.
  • Strong collaboration skills across infrastructure, research, and deployment teams.

Nice to Have

  • Experience training LLMs or other large transformer architectures.
  • Contributions to ML frameworks such as PyTorch, JAX, DeepSpeed, Megatron, or xFormers.
  • Familiarity with evaluation and serving frameworks such as vLLM and TensorRT-LLM.
  • Experience with data pipeline optimization, sharded datasets, or caching strategies.
  • Background in performance engineering, profiling, or low-level systems.
  • Papers at top-tier venues such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP.

Compensation & Benefits

  • Weekly lunch stipend of $75/£75 or equivalent in local currency.
  • Full health and dental benefits, including a separate mental-health budget.
  • RRSP matching, 401K, or pension scheme.
  • 100% parental-leave top-up for up to six months for either parent.
  • Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
  • Education and learning stipend for conferences, courses, and coaching.
  • Six weeks of paid vacation.
  • Travel budget for remote employees visiting other offices and an annual company offsite.
  • Co-working benefit and a $500 home-office stipend for remote employees.

Skills

JAX, PyTorch, Deepspeed, Megatron, Xformers, CUDA, Nccl, Fsdp, Zero, Kubernetes, Ray, Slurm, Docker, vLLM

Mercury

Mercury

San Francisco, CA
Senior Machine Learning Operations Engineer
$157k+/yrRemote5+ YOEML Engineering

Build and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.

Traba

Traba

New York, NY
Senior Software Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.

Front

Front

San Francisco, CA

Senior Applied AI Engineer
$205k+/yrHybrid5+ YOEML Engineering

Build and deploy generative AI and LLM-powered agentic applications at Front to automate customer support inquiries, enhance product capabilities, and drive operational insights. Requires 5+ years software engineering experience with strong production AI/ML focus, agentic/RAG expertise, and proficiency in Node.js, TS, and Python.

Baselayer

Baselayer

San Francisco, CA

Senior AI Engineer, Agentic Data Enrichment
$230k+/yrHybrid5+ YOEML Engineering

Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.

Blee

Blee

San Francisco, CA

Senior AI Engineer
$150k+/yrHybrid5+ YOEML Engineering

Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.