Machine Learning Infra Engineer
Builds and maintains ML training and inference infrastructure, scales distributed workloads across GPU clusters, and develops tooling for observability and fast iteration. Requires strong Python, systems engineering, Kubernetes, and distributed training experience.
About the job
What You’ll Do
- Build, and maintain our training and inference stack with an emphasis for fast iteration on training + flexibility for exploring new methods and high performance in inference.
- Develop benchmarks for both sets of stacks to identify bottlenecks.
- Explore SOTA advances in training and inference and work to apply them.
- Design systems for scaling model training across multi-node, multi-GPU environments with strong reliability and observability.
- Scale distributed training and inference workloads across large GPU clusters while improving utilization, reliability, and cost efficiency.
- Build the tooling, abstractions, and observability that help ML engineers move faster from experiment to production.
You’ll Thrive Here If You
- Hold yourself to a high bar for quality and precision.
- Enjoy solving complex problems and building from first principles.
- Have strong Python skills + a background in systems engineering.
- Are comfortable with Kubernetes and distributed training frameworks.
- Love getting your hands dirty with real-world implementation challenges.
- Operate well in fast-changing, high-growth environments.
- Collaborate effectively across technical and non-technical teams.
- Take full ownership from strategy through execution.
Bonus points if you
- Have experience at an early-stage or high-growth startup.
- Have developed in open source training/inference stacks in a meaningful way.
- Are excited to set up distributed inference across 100s-1000s of GPUs.
- Care deeply about combining technical excellence with business impact.
Benefits at Reducto
- Unlimited PTO
- Lunch: Receive a free lunch to eat with your teammates daily at the office
- Reimbursed Transportation
- Insurance: Generous health insurance covering medical, dental, and vision
- Health and Wellness Budget: Up to $150/mo reimbursement
- Parental Leave
Skills
Python, Kubernetes, Distributed Training, GPU, Multi-Node Training, Inference Frameworks, Training Frameworks, Observability, Benchmarks, Systems Engineering
Similar jobs
ML Engineering jobsBuild and operate production AI agents, automation workflows, and integrations that improve complex business processes. The role requires 5+ years of software engineering experience, modern LLM and agent-framework expertise, systems integration skills, and strong cross-functional collaboration.
Build and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.
Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.
Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.