Training: ML Framework Engineer
Develops and optimizes internal distributed ML training framework to boost hardware efficiency and enable researchers to experiment with new AI models. Requires strong Python skills, systems understanding, and passion for performance tuning.
About the job
Responsibilities
- Apply the latest techniques in our internal training framework to achieve impressive hardware efficiency for our training runs
- Profile and optimize our training framework
- Work with researchers to enable them to develop the next generation of models
Requirements
- Have run small scale ML experiments
- Love figuring out how systems work and continuously come up with ideas for how to make them faster while minimizing complexity and maintenance burden
- Have strong software engineering skills and are proficient in Python
Skills
Python, Distributed Systems, Machine Learning, PyTorch, TensorFlow, Gpu Programming, Performance Optimization, Profiling, Supercomputing
Similar jobs
ML Engineering jobsBuild and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.
Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.
Build and operate production machine-learning systems for search ranking, relevance, extraction quality, and LLM-driven features. The role requires production ML ownership, ranking or relevance expertise, large-scale data experience, Python, and rigorous experimentation skills.