Machine Learning Infrastructure Tech Lead
Lead ML infrastructure at Reducto by owning the training and inference stack. Hands-on role (80% building/optimizing) focused on GPU utilization, distributed systems, Kubernetes, kernels, and high-performance serving for AI document workflows. Requires 5+ years production ML infra experience and strong systems engineering skills.
About the job
Responsibilities
- Own the technical direction and roadmap for Reducto's ML infrastructure.
- Build and maintain our training and inference stack, balancing fast experimentation with high-performance production serving.
- Optimize model serving at every layer, including kernels, runtimes, batching, scheduling, and distributed inference.
- Design systems for reliable multi-node, multi-GPU training and inference.
- Improve GPU utilization, latency, throughput, reliability, observability, and cost efficiency.
- Develop benchmarks that identify bottlenecks and guide infrastructure investments.
- Evaluate state-of-the-art advances in training and inference and apply the ones that matter.
- Build the tooling and abstractions that help ML engineers move quickly from experiments to production.
- Partner with ML and Platform engineers on architecture, capacity planning, and technical prioritization.
- Raise the engineering bar through design reviews, mentorship, and hands-on technical leadership.
Requirements
- 5+ years of experience building production infrastructure, including significant ML systems experience.
- Led complex technical projects from an ambiguous problem through production deployment.
- Equally comfortable setting direction and personally implementing the hardest parts.
- Strong Python and systems-engineering skills.
- Understand the performance characteristics of modern GPU training or inference workloads.
- Comfortable with Kubernetes and distributed training or serving frameworks.
- Can reason across low-level model performance and higher-level platform architecture.
- Hold yourself to a high bar for quality, precision, and operational reliability.
- Operate well in a fast-changing, high-growth environment.
- Take full ownership from strategy through execution.
Nice-to-Haves
- Optimized or implemented CUDA, Triton, or custom model-serving kernels.
- Contributed meaningfully to frameworks such as vLLM, SGLang, PyTorch, TensorRT-LLM, Ray, or related open-source systems.
- Operated distributed inference or training across hundreds or thousands of GPUs.
- Built observability, scheduling, or capacity-management systems for GPU workloads.
- Experience at an early-stage or high-growth startup.
- Care deeply about connecting technical excellence to measurable business impact.
Skills
Python, Kubernetes, CUDA, Triton, PyTorch, Tensorrt-Llm, vLLM, Sglang, Ray, Gpu Optimization, Distributed Training, Model Serving
Similar jobs
ML Engineering jobsBuild and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.
Own machine learning end to end, from modeling messy clinical data through production deployment, monitoring, and infrastructure. The role requires 7+ years of experience building scalable ML systems, strong software and data engineering skills, and proficiency with Python, SQL, and cloud platforms.
Leads and manages an applied machine learning team developing production fraud detection and identity verification models. The role combines people leadership with hands-on technical work and requires substantial ML experience, production deployment expertise, and experience in risk-focused domains.
Builds and productionizes machine learning systems for trust and safety, including abuse detection, autonomous AI agents, and evaluation frameworks. The role requires 5+ years of applied ML experience, strong Python skills, experience with LLMs and scalable pipelines, and a relevant advanced degree or equivalent background.
Owns end-to-end production machine learning systems, including NLP, LLM, agentic, ranking, and recommendation capabilities. Requires 8+ years of industry experience, strong Python and cloud ML expertise, and the ability to deliver explainable AI products with cross-functional and customer impact.