Member Of Technical Staff, Post-Training
Build and post-train frontier AI models at scale, bridging research and production through scalable training software, distributed infrastructure, and performance optimization. The role requires strong software engineering skills and experience with large-model training and post-training.
About the job
Responsibilities
- Design and write high-performance, scalable software for training models.
- Post-train models to reach state-of-the-art performance.
- Coordinate with specialist teams, including Agentic and Code, to produce models with strong overall performance.
- Develop and implement techniques to improve training-cycle performance across supervised fine-tuning (SFT) and reinforcement learning (RL).
- Research, implement, and experiment with ideas using large-scale compute and data infrastructure.
- Contribute to production code and research efforts.
Requirements
- Extremely strong software engineering skills.
- Proficiency in Python and machine-learning frameworks including JAX, PyTorch, and XLA/MLIR.
- Experience with distributed training infrastructure such as Kubernetes and Slurm, and associated frameworks such as Ray.
- Experience using large-scale distributed training strategies.
- Hands-on experience training large models at scale.
- Hands-on experience with model post-training, with a strong emphasis on performance optimization.
Nice-to-haves
- Publications at top-tier venues such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP.
Compensation and Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Full health and dental benefits, including a separate mental-health budget.
- RRSP matching, 401(k), or pension scheme, depending on location.
- 100% parental-leave top-up for up to six months for either parent.
- Annual enrichment benefits covering arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation (30 working days).
- Travel budget for remote employees visiting other offices and an annual company offsite.
- Coworking benefit for employees not near an office.
- $500 home-office setup stipend.
Skills
Python, JAX, PyTorch, Xla, Mlir, Kubernetes, Slurm, Ray, Distributed Training, Supervised Fine-Tuning, Reinforcement Learning, Model Post-Training, Performance Optimization, LLMs
Similar jobs
ML Engineering jobsBuild and optimize a high-scale LLM inference engine spanning accelerator programming, host-device coordination, and distributed systems. The role requires strong systems programming, performance analysis, and an understanding of LLM inference across compute, memory, and interconnects.
Build and operate production AI agents, automation workflows, and integrations that improve complex business processes. The role requires 5+ years of software engineering experience, modern LLM and agent-framework expertise, systems integration skills, and strong cross-functional collaboration.
Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.