Machine Learning Performance Engineer - Offboard Training & Inference
Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.
215k – 285k/yr
On-siteML Engineering
About the role
Responsibilities
Profile and optimize distributed training end to end, including data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing.
Optimize large-scale offline and batch inference over petabyte-scale sensor logs through batching and scheduling strategies, quantization, low-precision execution, graph optimization, and accelerator saturation.
Establish roofline and performance models, quantify gaps between achieved and theoretical performance, and prioritize optimization opportunities by impact and effort.
Improve multi-node scaling efficiency through sharding and parallelism strategies, collective communication, interconnect utilization, memory-bandwidth optimization, and kernel fusion.
Drive cluster goodput by reducing GPU idle time caused by input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery.
Build benchmarking, observability, and regression-detection tooling.
Collaborate across engineering functions to solve complex data and compute problems at scale.
Contribute to a culture of collaboration, technical excellence, and innovation.
Requirements
Hands-on machine learning performance engineering experience, including profiling, roofline analysis, throughput optimization, and production root-cause investigation.
Experience with distributed multi-node training at scale, including FSDP, DeepSpeed, Megatron, NCCL, or equivalent, and diagnosing scaling inefficiency as node count grows.
Deep familiarity with GPU or accelerator performance concepts, including memory bandwidth, kernel launch overhead, occupancy, quantization, and collective communication.
Experience with high-throughput or batch inference systems such as NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar.
Fluency in Python and proficiency in C++ or another systems language.
Excellent debugging, analytical, and problem-solving skills.
Deep understanding of machine learning foundations and the ability to develop technical solutions for problems without an established playbook.
Nice to Have
GPU kernel development experience with CUDA, Triton, CUTLASS, or hand-tuned attention implementations.
Experience with profiling toolchains such as Nsight Systems, Nsight Compute, PyTorch Profiler, or perf.
Experience with GPU scheduling and orchestration on Kubernetes, Slurm, or Ray, including multi-tenant cluster utilization.
Experience with fault tolerance and elastic training for long-running jobs, including checkpointing strategy, straggler mitigation, and preemption recovery.
Familiarity with autonomy or robotics data, including ROS, OpenCV, and multi-sensor log formats.
Build research and production infrastructure for replayable enterprise environments, agent evaluation, and model improvement. The role requires 4+ years building production software or ML systems, including 2+ years in reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or related work.
215k – 330k/yrOn-site4+ YOEML Engineering
Machine Learning Engineer
LiftoffCalifornia
Machine Learning Engineer building statistical models, optimization systems, and experiments for mobile ad tech economics on the Revenue Engine team. Requires PhD in CS/ML/Economics and industry experience applying ML or economics at scale.
215k – 275k/yrRemoteML Engineering
AI Research Engineer
HexSan Francisco, CA +1
Builds and deploys production AI features like Notebook Agent for data science workflows, partnering with product teams on experiments, model fine-tuning, and infra. Requires senior AI/ML engineering experience with MLOps, Python/TS proficiency.
214k – 285k/yrHybrid5+ YOEML Engineering
Software Engineer, Enterprise AI
Scale AINew York, NY +1
Build and scale enterprise Generative AI platform, owning large product areas across backend, frontend, LLMs, and ML models. Requires 4+ years experience, proficiency in Python/JavaScript/SQL, Kubernetes, and cloud providers.
216k – 270k/yrOn-site4+ YOEML Engineering
ML Research Engineer, ML Systems
Scale AISan Francisco, CA +2
Builds and optimizes distributed frameworks for LLM training and inference on Scale's RLXF platform. Collaborates with ML teams to accelerate research, requiring expertise in PyTorch, CUDA, transformers, and large-scale distributed systems.