Skip to content
ReductoReductoSan Francisco, CA

Machine Learning Infrastructure Tech Lead

Lead ML infrastructure at Reducto by owning the training and inference stack. Hands-on role (80% building/optimizing) focused on GPU utilization, distributed systems, Kubernetes, kernels, and high-performance serving for AI document workflows. Requires 5+ years production ML infra experience and strong systems engineering skills.

200k – 300k/yr
On-site7+ YOEML Engineering

About the role

Responsibilities

  • Own the technical direction and roadmap for Reducto's ML infrastructure.
  • Build and maintain our training and inference stack, balancing fast experimentation with high-performance production serving.
  • Optimize model serving at every layer, including kernels, runtimes, batching, scheduling, and distributed inference.
  • Design systems for reliable multi-node, multi-GPU training and inference.
  • Improve GPU utilization, latency, throughput, reliability, observability, and cost efficiency.
  • Develop benchmarks that identify bottlenecks and guide infrastructure investments.
  • Evaluate state-of-the-art advances in training and inference and apply the ones that matter.
  • Build the tooling and abstractions that help ML engineers move quickly from experiments to production.
  • Partner with ML and Platform engineers on architecture, capacity planning, and technical prioritization.
  • Raise the engineering bar through design reviews, mentorship, and hands-on technical leadership.

Requirements

  • 5+ years of experience building production infrastructure, including significant ML systems experience.
  • Led complex technical projects from an ambiguous problem through production deployment.
  • Equally comfortable setting direction and personally implementing the hardest parts.
  • Strong Python and systems-engineering skills.
  • Understand the performance characteristics of modern GPU training or inference workloads.
  • Comfortable with Kubernetes and distributed training or serving frameworks.
  • Can reason across low-level model performance and higher-level platform architecture.
  • Hold yourself to a high bar for quality, precision, and operational reliability.
  • Operate well in a fast-changing, high-growth environment.
  • Take full ownership from strategy through execution.

Nice-to-Haves

  • Optimized or implemented CUDA, Triton, or custom model-serving kernels.
  • Contributed meaningfully to frameworks such as vLLM, SGLang, PyTorch, TensorRT-LLM, Ray, or related open-source systems.
  • Operated distributed inference or training across hundreds or thousands of GPUs.
  • Built observability, scheduling, or capacity-management systems for GPU workloads.
  • Experience at an early-stage or high-growth startup.
  • Care deeply about connecting technical excellence to measurable business impact.

Skills

PythonKubernetesCUDAtritonPyTorchtensorrt-llmvLLMsglangRaygpu optimizationDistributed Trainingmodel serving

Similar roles

ML Engineering jobs
Traba

Senior Software Engineer

TrabaNew York, NY +1

Build and own production AI agent systems (harnesses, evals, orchestration) on frontier LLMs for industrial supply chain workflows at Traba. Requires 5+ years software engineering with 1+ year shipping LLM/agent features, strong Python/TS, and high-agency in ambiguous customer environments.

200k – 240k/yr
Hybrid5+ YOEML Engineering
Rad AI

Machine Learning Research Manager

Rad AISan Francisco, CA

Lead and mentor a team of applied and clinical researchers as a player-coach. Guide ML research in NLP, LLMs, and clinical applications for radiology, translating ideas into production systems while partnering with clinicians and engineers. Requires MS/PhD and 6+ years applied ML research experience.

200k – 230k/yr
On-site7+ YOEML Engineering
Airbnb

Senior Machine Learning Engineer, Relevance and Personalization

AirbnbUnited States

Build and productionize cutting-edge ML models for Airbnb's query intelligence, including autocomplete, query tagging, expansion, intent modeling, and LLM-powered natural language search to understand guest intent.

200k – 235k/yr
Remote5+ YOEML Engineering
Roger Healthcare

Senior Applied AI Engineer

Roger HealthcareSan Francisco, CA

Senior Applied AI Engineer building the core intelligence layer for Roger, an AI platform for home health clinicians. Responsibilities include training/fine-tuning LLMs on proprietary clinical data, building rigorous eval and monitoring systems, and shipping reliable agentic LLM workflows that improve patient care.

200k – 250k/yr
On-site7+ YOEML Engineering
Astronomer

Senior Software Engineer, Build

AstronomerNew York, NY

Build and scale Astronomer's AI-powered global context layer for data, focusing on semantic search, retrieval, code generation, and applied AI for data engineering workflows. Requires 5+ years software engineering experience with Python or Go, plus strong interest in LLMs and data tools.

200k – 230k/yr
Hybrid5+ YOEML Engineering