Senior ML Systems Engineer, Frameworks & Tooling
Build and evolve the distributed training framework and tooling powering frontier-scale language models. The role focuses on large-scale ML systems, HPC infrastructure, performance optimization, reliability, and developer tooling across multi-node GPU clusters.
About the job
Responsibilities
- Build and own the training framework responsible for large-scale LLM training.
- Design distributed training abstractions, including data, tensor, and pipeline parallelism, FSDP/ZeRO strategies, memory management, and checkpointing.
- Improve training throughput and stability on multi-node clusters, including GB200/300, AMD, and H200/100 systems.
- Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
- Collaborate with infrastructure teams to ensure cluster, container, and hardware configurations support high-performance training.
- Investigate and resolve performance bottlenecks across the ML systems stack.
- Build robust systems for reproducible, debuggable, large-scale runs.
- Build high-performance data loading and caching pipelines.
- Implement performance profiling across the ML systems stack.
- Develop internal metrics and monitoring for training runs.
- Build reproducibility and regression-testing infrastructure.
- Develop performant, fault-tolerant distributed checkpointing systems.
Requirements
- Strong engineering experience in large-scale distributed training or HPC systems.
- Deep familiarity with JAX internals, distributed training libraries, or custom kernels and fused operations.
- Experience with multi-node cluster orchestration.
- Comfort debugging performance issues across CUDA/NCCL, networking, I/O, and data pipelines.
- Experience with containerized environments.
- Track record of building tools that increase developer velocity for ML teams.
- Excellent judgment regarding performance versus complexity and research velocity versus maintainability.
- Strong collaboration skills across infrastructure, research, and deployment teams.
Nice to Have
- Experience training LLMs or other large transformer architectures.
- Contributions to ML frameworks such as PyTorch, JAX, DeepSpeed, Megatron, or xFormers.
- Familiarity with evaluation and serving frameworks such as vLLM and TensorRT-LLM.
- Experience with data pipeline optimization, sharded datasets, or caching strategies.
- Background in performance engineering, profiling, or low-level systems.
- Papers at top-tier venues such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP.
Compensation & Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Full health and dental benefits, including a separate mental-health budget.
- RRSP matching, 401K, or pension scheme.
- 100% parental-leave top-up for up to six months for either parent.
- Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation.
- Travel budget for remote employees visiting other offices and an annual company offsite.
- Co-working benefit and a $500 home-office stipend for remote employees.
Skills
JAX, PyTorch, Deepspeed, Megatron, Xformers, CUDA, Nccl, Fsdp, Zero, Kubernetes, Ray, Slurm, Docker, vLLM
Similar jobs
ML Engineering jobsBuild and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.
Build and deploy generative AI and LLM-powered agentic applications at Front to automate customer support inquiries, enhance product capabilities, and drive operational insights. Requires 5+ years software engineering experience with strong production AI/ML focus, agentic/RAG expertise, and proficiency in Node.js, TS, and Python.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.
Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.