Member of Technical Staff, Training Performance Engineer
Optimizes training performance for advanced language models by developing scalable software, GPU kernels, distributed training systems, and profiling tools. The role requires strong software engineering skills, Python and ML framework proficiency, and experience with CUDA or Triton.
Salary not listed
HybridML Engineering
About the role
Responsibilities
Design and write high-performance, scalable software for model training.
Understand architectural modifications and design choices and their effects on training throughput and quality.
Write low-level CUDA and Triton kernels to maximize accelerator performance.
Research, implement, and experiment with ideas on supercomputing and data infrastructure.
Identify and remove performance bottlenecks.
Develop training and profiling tools to improve model performance.
Collaborate with researchers and engineers on large-scale language model systems.
Requirements
Extremely strong software engineering skills.
Proficiency in Python and related machine learning frameworks, including JAX, PyTorch, and XLA/MLIR.
Experience writing GPU kernels using CUDA, Triton, or similar technologies.
Experience with large-scale distributed training strategies.
Familiarity with autoregressive sequence models such as Transformers.
Nice to Have
A paper published at a top-tier venue such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP.
Compensation and Benefits
Weekly lunch stipend of $75/£75 or equivalent in local currency.
Full health and dental benefits, including a separate mental health budget.
RRSP matching, 401(k), and pension scheme.
100% parental leave top-up for up to six months for either parent.
Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
Education and learning stipend for conferences, courses, and coaching.
Six weeks of paid vacation (30 working days).
Travel budget for remote employees visiting other offices and an annual company offsite.
Coworking benefit for employees not near an office.
Architects and operates production machine-learning systems that classify web and API traffic, detect bots and scrapers, and support real-time mitigation at internet edge latency. The role requires 9+ years of applied ML experience in adversarial domains and strong expertise in evaluation, data pipelines, and large-scale systems.
212k – 265k/yrRemote9+ YOEML Engineering
Member of Technical Staff, Agentic Environments
CohereNew York, NY
Build scalable software and tools for frontier model training, research experimentation, and production machine learning systems. The role requires strong Python and distributed-training expertise, experience with ML frameworks and infrastructure, and the ability to optimize and debug large language model systems.
Leads the technical direction of large-scale ML infrastructure for embedding, recommendation, and personalization systems. The role requires 8+ years of ML engineering experience, expertise in deep learning and distributed training, and strong leadership across research, infrastructure, and production deployment.
Leads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.
266k – 372k/yrRemote10+ YOEML Engineering
Staff Applied Scientist
Garner HealthNew York, NY
Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.