Skip to content
Applied IntuitionApplied IntuitionSunnyvale, CA

Machine Learning Performance Engineer - Offboard Training & Inference

Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.

215k – 285k/yr
On-siteML Engineering

About the role

Responsibilities

  • Profile and optimize distributed training end to end, including data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing.
  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs through batching and scheduling strategies, quantization, low-precision execution, graph optimization, and accelerator saturation.
  • Establish roofline and performance models, quantify gaps between achieved and theoretical performance, and prioritize optimization opportunities by impact and effort.
  • Improve multi-node scaling efficiency through sharding and parallelism strategies, collective communication, interconnect utilization, memory-bandwidth optimization, and kernel fusion.
  • Drive cluster goodput by reducing GPU idle time caused by input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery.
  • Build benchmarking, observability, and regression-detection tooling.
  • Collaborate across engineering functions to solve complex data and compute problems at scale.
  • Contribute to a culture of collaboration, technical excellence, and innovation.

Requirements

  • Hands-on machine learning performance engineering experience, including profiling, roofline analysis, throughput optimization, and production root-cause investigation.
  • Experience with distributed multi-node training at scale, including FSDP, DeepSpeed, Megatron, NCCL, or equivalent, and diagnosing scaling inefficiency as node count grows.
  • Deep familiarity with GPU or accelerator performance concepts, including memory bandwidth, kernel launch overhead, occupancy, quantization, and collective communication.
  • Experience with high-throughput or batch inference systems such as NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar.
  • Fluency in Python and proficiency in C++ or another systems language.
  • Excellent debugging, analytical, and problem-solving skills.
  • Deep understanding of machine learning foundations and the ability to develop technical solutions for problems without an established playbook.

Nice to Have

  • GPU kernel development experience with CUDA, Triton, CUTLASS, or hand-tuned attention implementations.
  • Experience with profiling toolchains such as Nsight Systems, Nsight Compute, PyTorch Profiler, or perf.
  • Experience with GPU scheduling and orchestration on Kubernetes, Slurm, or Ray, including multi-tenant cluster utilization.
  • Experience with fault tolerance and elastic training for long-running jobs, including checkpointing strategy, straggler mitigation, and preemption recovery.
  • Familiarity with autonomy or robotics data, including ROS, OpenCV, and multi-sensor log formats.

Skills

PythonC++fsdpdeepspeedmegatronncclCUDAtritonTensorRTonnx runtimeRayKubernetesslurmnsight systemspytorch profiler

Similar roles

ML Engineering jobs
Ambral

Founding Research Engineer

AmbralNew York, NY +1

Build research and production infrastructure for replayable enterprise environments, agent evaluation, and model improvement. The role requires 4+ years building production software or ML systems, including 2+ years in reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or related work.

215k – 330k/yrOn-site4+ YOEML Engineering
Liftoff

Machine Learning Engineer

LiftoffCalifornia

Machine Learning Engineer building statistical models, optimization systems, and experiments for mobile ad tech economics on the Revenue Engine team. Requires PhD in CS/ML/Economics and industry experience applying ML or economics at scale.

215k – 275k/yrRemoteML Engineering
Hex

AI Research Engineer

HexSan Francisco, CA +1

Builds and deploys production AI features like Notebook Agent for data science workflows, partnering with product teams on experiments, model fine-tuning, and infra. Requires senior AI/ML engineering experience with MLOps, Python/TS proficiency.

214k – 285k/yrHybrid5+ YOEML Engineering
Scale AI

Software Engineer, Enterprise AI

Scale AINew York, NY +1

Build and scale enterprise Generative AI platform, owning large product areas across backend, frontend, LLMs, and ML models. Requires 4+ years experience, proficiency in Python/JavaScript/SQL, Kubernetes, and cloud providers.

216k – 270k/yrOn-site4+ YOEML Engineering
Scale AI

ML Research Engineer, ML Systems

Scale AISan Francisco, CA +2

Builds and optimizes distributed frameworks for LLM training and inference on Scale's RLXF platform. Collaborates with ML teams to accelerate research, requiring expertise in PyTorch, CUDA, transformers, and large-scale distributed systems.

218k – 273k/yrOn-siteML Engineering