Skip to content
ReductoReducto

Machine Learning Infrastructure Tech Lead

Lead ML infrastructure at Reducto by owning the training and inference stack. Hands-on role (80% building/optimizing) focused on GPU utilization, distributed systems, Kubernetes, kernels, and high-performance serving for AI document workflows. Requires 5+ years production ML infra experience and strong systems engineering skills.

About the job

Responsibilities

  • Own the technical direction and roadmap for Reducto's ML infrastructure.
  • Build and maintain our training and inference stack, balancing fast experimentation with high-performance production serving.
  • Optimize model serving at every layer, including kernels, runtimes, batching, scheduling, and distributed inference.
  • Design systems for reliable multi-node, multi-GPU training and inference.
  • Improve GPU utilization, latency, throughput, reliability, observability, and cost efficiency.
  • Develop benchmarks that identify bottlenecks and guide infrastructure investments.
  • Evaluate state-of-the-art advances in training and inference and apply the ones that matter.
  • Build the tooling and abstractions that help ML engineers move quickly from experiments to production.
  • Partner with ML and Platform engineers on architecture, capacity planning, and technical prioritization.
  • Raise the engineering bar through design reviews, mentorship, and hands-on technical leadership.

Requirements

  • 5+ years of experience building production infrastructure, including significant ML systems experience.
  • Led complex technical projects from an ambiguous problem through production deployment.
  • Equally comfortable setting direction and personally implementing the hardest parts.
  • Strong Python and systems-engineering skills.
  • Understand the performance characteristics of modern GPU training or inference workloads.
  • Comfortable with Kubernetes and distributed training or serving frameworks.
  • Can reason across low-level model performance and higher-level platform architecture.
  • Hold yourself to a high bar for quality, precision, and operational reliability.
  • Operate well in a fast-changing, high-growth environment.
  • Take full ownership from strategy through execution.

Nice-to-Haves

  • Optimized or implemented CUDA, Triton, or custom model-serving kernels.
  • Contributed meaningfully to frameworks such as vLLM, SGLang, PyTorch, TensorRT-LLM, Ray, or related open-source systems.
  • Operated distributed inference or training across hundreds or thousands of GPUs.
  • Built observability, scheduling, or capacity-management systems for GPU workloads.
  • Experience at an early-stage or high-growth startup.
  • Care deeply about connecting technical excellence to measurable business impact.

Skills

Python, Kubernetes, CUDA, Triton, PyTorch, Tensorrt-Llm, vLLM, Sglang, Ray, Gpu Optimization, Distributed Training, Model Serving

Traba

Traba

New York, NY
Senior Software Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.

Metriport

Metriport

San Francisco, CA

Senior AI/ML Engineer
$200k+/yrHybrid7+ YOEML Engineering

Own machine learning end to end, from modeling messy clinical data through production deployment, monitoring, and infrastructure. The role requires 7+ years of experience building scalable ML systems, strong software and data engineering skills, and proficiency with Python, SQL, and cloud platforms.

SentiLink

SentiLink

United States

Applied Machine Learning Manager - Application Fraud
$200k+/yrRemote6+ YOEML Engineering

Leads and manages an applied machine learning team developing production fraud detection and identity verification models. The role combines people leadership with hands-on technical work and requires substantial ML experience, production deployment expertise, and experience in risk-focused domains.

Airbnb

Airbnb

United States

Senior Machine Learning Engineer, Trust
$200k+/yrRemote5+ YOEML Engineering

Builds and productionizes machine learning systems for trust and safety, including abuse detection, autonomous AI agents, and evaluation frameworks. The role requires 5+ years of applied ML experience, strong Python skills, experience with LLMs and scalable pipelines, and a relevant advanced degree or equivalent background.

6sense

6sense

San Francisco, CA

Senior Machine Learning Engineer
$200k+/yrRemote8+ YOEML Engineering

Owns end-to-end production machine learning systems, including NLP, LLM, agentic, ranking, and recommendation capabilities. Requires 8+ years of industry experience, strong Python and cloud ML expertise, and the ability to deliver explainable AI products with cross-functional and customer impact.