# Machine Learning Performance Engineer - Offboard Training & Inference

**Company:** [Applied Intuition](https://hotfix.jobs/companies/applied-intuition)
**Location:** Sunnyvale, CA
**Role:** ML Engineering
**Salary:** $215k – $285k/yr
**Skills:** Python, C++, fsdp, deepspeed, megatron, nccl, CUDA, triton, TensorRT, onnx runtime, Ray, Kubernetes, slurm, nsight systems, pytorch profiler
**Posted:** 2026-08-12

> Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.

## Job Description

## Responsibilities
- Profile and optimize distributed training end to end, including data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing.
- Optimize large-scale offline and batch inference over petabyte-scale sensor logs through batching and scheduling strategies, quantization, low-precision execution, graph optimization, and accelerator saturation.
- Establish roofline and performance models, quantify gaps between achieved and theoretical performance, and prioritize optimization opportunities by impact and effort.
- Improve multi-node scaling efficiency through sharding and parallelism strategies, collective communication, interconnect utilization, memory-bandwidth optimization, and kernel fusion.
- Drive cluster goodput by reducing GPU idle time caused by input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery.
- Build benchmarking, observability, and regression-detection tooling.
- Collaborate across engineering functions to solve complex data and compute problems at scale.
- Contribute to a culture of collaboration, technical excellence, and innovation.

## Requirements
- Hands-on machine learning performance engineering experience, including profiling, roofline analysis, throughput optimization, and production root-cause investigation.
- Experience with distributed multi-node training at scale, including FSDP, DeepSpeed, Megatron, NCCL, or equivalent, and diagnosing scaling inefficiency as node count grows.
- Deep familiarity with GPU or accelerator performance concepts, including memory bandwidth, kernel launch overhead, occupancy, quantization, and collective communication.
- Experience with high-throughput or batch inference systems such as NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar.
- Fluency in Python and proficiency in C++ or another systems language.
- Excellent debugging, analytical, and problem-solving skills.
- Deep understanding of machine learning foundations and the ability to develop technical solutions for problems without an established playbook.

## Nice to Have
- GPU kernel development experience with CUDA, Triton, CUTLASS, or hand-tuned attention implementations.
- Experience with profiling toolchains such as Nsight Systems, Nsight Compute, PyTorch Profiler, or perf.
- Experience with GPU scheduling and orchestration on Kubernetes, Slurm, or Ray, including multi-tenant cluster utilization.
- Experience with fault tolerance and elastic training for long-running jobs, including checkpointing strategy, straggler mitigation, and preemption recovery.
- Familiarity with autonomy or robotics data, including ROS, OpenCV, and multi-sensor log formats.

## Similar roles

- [Founding Research Engineer](https://hotfix.jobs/jobs/92328d5a-8c83-43a9-afb7-e29b5480350f) - Ambral - New York, NY - $215k – $330k/yr
- [Machine Learning Engineer](https://hotfix.jobs/jobs/97713038-58cb-46c5-a9c7-0e1554e1ba4f) - Liftoff - Remote - $215k – $275k/yr
- [AI Research Engineer](https://hotfix.jobs/jobs/eb72cd41-2cc0-499f-b1aa-f5721d12433c) - Hex - San Francisco, CA - $214k – $285k/yr
- [Software Engineer, Enterprise AI](https://hotfix.jobs/jobs/98385fe7-6f4c-4386-ab6e-d4033a6d2440) - Scale AI - New York, NY - $216k – $270k/yr
- [ML Research Engineer, ML Systems](https://hotfix.jobs/jobs/c814c5a3-7e79-4c03-8f4b-89eb8352da30) - Scale AI - San Francisco, CA - $218k – $273k/yr

**Apply:** https://hotfix.jobs/jobs/b2f397d3-cd2d-4eeb-bc02-1dfe0c90a5a7
**Canonical:** https://hotfix.jobs/jobs/b2f397d3-cd2d-4eeb-bc02-1dfe0c90a5a7