ML Research Engineer, ML Systems
Builds and optimizes distributed frameworks for LLM training and inference on Scale's RLXF platform. Collaborates with ML teams to accelerate research, requiring expertise in PyTorch, CUDA, transformers, and large-scale distributed systems.
About the job
You will:
- Build, profile and optimize our training and inference framework
- Collaborate with ML teams to accelerate their research and development and enable them to develop the next generation of models and data curation
- Research and integrate state-of-the-art technologies to optimize our ML system
Ideally you’d have:
- Strong excitement about system optimization
- Experience with multi-node LLM training and inference
- Experience with developing large-scale distributed ML systems
- Strong software engineering skills, proficient in frameworks and tools such as CUDA, PyTorch, transformers, flash attention, etc.
- Strong written and verbal communication skills and the ability to operate in a cross functional team environment
Nice to haves:
- Demonstrated expertise in post-training methods or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc.
Skills
PyTorch, CUDA, Transformers, Flash Attention, Llm Training, Distributed Ml Systems, RLHF, Multi-Node Training, Ml Inference, System Optimization
Similar jobs
ML Engineering jobsBuild and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.
Build AI agent harnesses, models, and product capabilities that enable agents to perform complex work across digital environments. The role combines applied AI research and software engineering, requiring Python proficiency, strong product judgment, and experience with agent tooling, reinforcement learning, or browser technologies.
Builds the platform, verifiers, environments, and grading infrastructure used to evaluate enterprise AI agents at scale. The role combines strong software engineering with expertise in agent runtimes, evaluation design, benchmarks, and production failure analysis.
Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.
Build and operate production machine-learning systems for search ranking, relevance, extraction quality, and LLM-driven features. The role requires production ML ownership, ranking or relevance expertise, large-scale data experience, Python, and rigorous experimentation skills.