# Machine Learning Infrastructure Tech Lead

**Company:** [Reducto](https://hotfix.jobs/companies/reducto)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $200k – $300k/yr
**Experience:** 7+ years
**Skills:** Python, Kubernetes, CUDA, triton, PyTorch, tensorrt-llm, vLLM, sglang, Ray, gpu optimization, Distributed Training, model serving
**Posted:** 2026-07-16

> Lead ML infrastructure at Reducto by owning the training and inference stack. Hands-on role (80% building/optimizing) focused on GPU utilization, distributed systems, Kubernetes, kernels, and high-performance serving for AI document workflows. Requires 5+ years production ML infra experience and strong systems engineering skills.

## Job Description

## Responsibilities
- Own the technical direction and roadmap for Reducto's ML infrastructure.
- Build and maintain our training and inference stack, balancing fast experimentation with high-performance production serving.
- Optimize model serving at every layer, including kernels, runtimes, batching, scheduling, and distributed inference.
- Design systems for reliable multi-node, multi-GPU training and inference.
- Improve GPU utilization, latency, throughput, reliability, observability, and cost efficiency.
- Develop benchmarks that identify bottlenecks and guide infrastructure investments.
- Evaluate state-of-the-art advances in training and inference and apply the ones that matter.
- Build the tooling and abstractions that help ML engineers move quickly from experiments to production.
- Partner with ML and Platform engineers on architecture, capacity planning, and technical prioritization.
- Raise the engineering bar through design reviews, mentorship, and hands-on technical leadership.

## Requirements
- 5+ years of experience building production infrastructure, including significant ML systems experience.
- Led complex technical projects from an ambiguous problem through production deployment.
- Equally comfortable setting direction and personally implementing the hardest parts.
- Strong Python and systems-engineering skills.
- Understand the performance characteristics of modern GPU training or inference workloads.
- Comfortable with Kubernetes and distributed training or serving frameworks.
- Can reason across low-level model performance and higher-level platform architecture.
- Hold yourself to a high bar for quality, precision, and operational reliability.
- Operate well in a fast-changing, high-growth environment.
- Take full ownership from strategy through execution.

## Nice-to-Haves
- Optimized or implemented CUDA, Triton, or custom model-serving kernels.
- Contributed meaningfully to frameworks such as vLLM, SGLang, PyTorch, TensorRT-LLM, Ray, or related open-source systems.
- Operated distributed inference or training across hundreds or thousands of GPUs.
- Built observability, scheduling, or capacity-management systems for GPU workloads.
- Experience at an early-stage or high-growth startup.
- Care deeply about connecting technical excellence to measurable business impact.

## Similar roles

- [Senior Software Engineer](https://hotfix.jobs/jobs/17e0432e-a11f-4a57-8de2-0c603c7b5780) - Traba - New York, NY - $200k – $240k/yr
- [Machine Learning Research Manager](https://hotfix.jobs/jobs/d6047c03-2cd4-40fd-92a5-332a272510cc) - Rad AI - San Francisco, CA - $200k – $230k/yr
- [Senior Machine Learning Engineer, Relevance and Personalization](https://hotfix.jobs/jobs/589feee3-05a3-4944-84df-8bd20d220687) - Airbnb - Remote - $200k – $235k/yr
- [Senior Applied AI Engineer](https://hotfix.jobs/jobs/2ee355ea-77df-4a9a-bab1-4a96c2ed3514) - Roger Healthcare - San Francisco, CA - $200k – $250k/yr
- [Senior Software Engineer, Build](https://hotfix.jobs/jobs/f9da1ebf-fa44-456b-84fa-e08ad0706de2) - Astronomer - New York, NY - $200k – $230k/yr

**Apply:** https://hotfix.jobs/jobs/c794b0f7-3284-4f6b-9529-276ba3dae6f6
**Canonical:** https://hotfix.jobs/jobs/c794b0f7-3284-4f6b-9529-276ba3dae6f6