Skip to content
ReductoReducto

Machine Learning Infra Engineer

Builds and maintains ML training and inference infrastructure, scales distributed workloads across GPU clusters, and develops tooling for observability and fast iteration. Requires strong Python, systems engineering, Kubernetes, and distributed training experience.

About the job

What You’ll Do

  • Build, and maintain our training and inference stack with an emphasis for fast iteration on training + flexibility for exploring new methods and high performance in inference.
  • Develop benchmarks for both sets of stacks to identify bottlenecks.
  • Explore SOTA advances in training and inference and work to apply them.
  • Design systems for scaling model training across multi-node, multi-GPU environments with strong reliability and observability.
  • Scale distributed training and inference workloads across large GPU clusters while improving utilization, reliability, and cost efficiency.
  • Build the tooling, abstractions, and observability that help ML engineers move faster from experiment to production.

You’ll Thrive Here If You

  • Hold yourself to a high bar for quality and precision.
  • Enjoy solving complex problems and building from first principles.
  • Have strong Python skills + a background in systems engineering.
  • Are comfortable with Kubernetes and distributed training frameworks.
  • Love getting your hands dirty with real-world implementation challenges.
  • Operate well in fast-changing, high-growth environments.
  • Collaborate effectively across technical and non-technical teams.
  • Take full ownership from strategy through execution.

Bonus points if you

  • Have experience at an early-stage or high-growth startup.
  • Have developed in open source training/inference stacks in a meaningful way.
  • Are excited to set up distributed inference across 100s-1000s of GPUs.
  • Care deeply about combining technical excellence with business impact.

Benefits at Reducto

  • Unlimited PTO
  • Lunch: Receive a free lunch to eat with your teammates daily at the office
  • Reimbursed Transportation
  • Insurance: Generous health insurance covering medical, dental, and vision
  • Health and Wellness Budget: Up to $150/mo reimbursement
  • Parental Leave

Skills

Python, Kubernetes, Distributed Training, GPU, Multi-Node Training, Inference Frameworks, Training Frameworks, Observability, Benchmarks, Systems Engineering

OnePay

OnePay

United States

Forward Deployed Engineer (AI and Automation)
$150k+/yrRemote5+ YOEML Engineering

Build and operate production AI agents, automation workflows, and integrations that improve complex business processes. The role requires 5+ years of software engineering experience, modern LLM and agent-framework expertise, systems integration skills, and strong cross-functional collaboration.

LangChain

LangChain

New York, NY
AI Engineer, Enablement
$150k+/yrOn-site3+ YOEML Engineering

Build and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.

Roboflow

Roboflow

San Francisco, CA

Member of Technical Staff — Frontier Data
$150k+/yrRemoteML Engineering

Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.

Beacon Biosignals

Beacon Biosignals

Boston, MA
Algorithm Engineer
$150k+/yrRemote4+ YOEML Engineering

Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Software Engineer - Prediction and Planning ML
$151k+/yrOn-site3+ YOEML Engineering

Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.