Build scalable software and tools for frontier model training, research experimentation, and production machine learning systems. The role requires strong Python and distributed-training expertise, experience with ML frameworks and infrastructure, and the ability to optimize and debug large language model systems.
Salary not listed
RemoteML Engineering
About the role
Responsibilities
Design and write high-performing, scalable software for training models.
Develop tools to support and accelerate research and large language model training.
Coordinate with Infrastructure, Efficiency, Serving, Agent, Multimodal, and Multilingual teams.
Research, implement, and experiment with ideas on cluster and data infrastructure.
Collaborate with scientists, engineers, and cross-functional teams.
Requirements
Extremely strong software engineering skills.
Commitment to test-driven development, clean code, and reducing technical debt.
Proficiency in Python and machine learning frameworks such as JAX, PyTorch, and/or XLA/MLIR.
Experience using and debugging large-scale distributed training strategies, including memory and speed profiling.
Experience with distributed training infrastructure such as Kubernetes and associated frameworks such as Ray.
Experience in machine learning and large language model academic research.
Ability to tune and optimize large language models.
Ability to navigate complex machine learning codebases and resolve issues.
Ability to work effectively with colleagues across experience levels in fast-paced, technically challenging environments.
Compensation and Benefits
Weekly lunch stipend of $75/£75 or equivalent in local currency.
Full health and dental benefits, including a separate mental health budget.
RRSP matching, 401(k), or pension scheme.
100% parental leave top-up for up to six months for either parent.
Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
Education and learning stipend for conferences, courses, and coaching.
Six weeks of paid vacation (30 working days).
Travel budget for remote employees visiting other offices and an annual company offsite.
Coworking benefit for employees not near an office.
$500 home office stipend.
Skills
PythonJAXPyTorchxlamlirKubernetesRayDistributed Trainingmemory profilingspeed profilingLLMsMachine Learningtest-driven development
Architects and operates production machine-learning systems that classify web and API traffic, detect bots and scrapers, and support real-time mitigation at internet edge latency. The role requires 9+ years of applied ML experience in adversarial domains and strong expertise in evaluation, data pipelines, and large-scale systems.
Leads the technical direction of large-scale ML infrastructure for embedding, recommendation, and personalization systems. The role requires 8+ years of ML engineering experience, expertise in deep learning and distributed training, and strong leadership across research, infrastructure, and production deployment.
Leads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.
266k – 372k/yrRemote10+ YOEML Engineering
Staff Applied Scientist
Garner HealthNew York, NY
Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.
300k – 390k/yrHybrid7+ YOEML Engineering
Senior/Staff Software Engineer, ML Inference Platform
NuroMountain View, CA
Build and operate ML infrastructure for autonomy teams, including training and deployment pipelines, model observability, inference serving, and compiler platforms across hardware targets. Requires a degree, 3+ years of relevant experience, Python proficiency, and distributed-systems expertise.