Skip to content

Machine Learning Infrastructure Engineer

Designs, builds, and maintains ML training and serving infrastructure, providing support to research teams. Requires 4+ years in ML infrastructure, cloud platforms like Kubernetes and Google Cloud, and GPU experience.

About the job

Responsibilities

  • Provide infrastructure support to our ML research and product
  • Build tooling to diagnose cluster issues and hardware failures
  • Monitor deployments, manage experiments, and generally support our research
  • Maximize GPU allocation and utilization for both serving and training

Requirements

  • 4+ years of experience supporting the infrastructure within an ML environment
  • Experience in developing tools used to diagnose ML infrastructure problems and failures
  • Experience with cloud platforms (e.g., Compute Engine, Kubernetes, Cloud Storage)
  • Experience working with GPUs

Nice to Have

  • Experience with large GPU clusters and high-performance computing/networking
  • Experience with supporting large language model training
  • Experience with ML frameworks like Pytorch/TensorFlow/JAX
  • Experience with GPU kernel development

Skills

Kubernetes, GCP, Compute Engine, Cloud Storage, Gpus, PyTorch, TensorFlow, JAX, Gpu Kernel Development, Large Gpu Clusters

LangChain

LangChain

New York, NY
AI Engineer, Enablement
$150k+/yrOn-site3+ YOEML Engineering

Build and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.

Roboflow

Roboflow

San Francisco, CA

Member of Technical Staff — Frontier Data
$150k+/yrRemoteML Engineering

Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.

Beacon Biosignals

Beacon Biosignals

Boston, MA
Algorithm Engineer
$150k+/yrRemote4+ YOEML Engineering

Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Software Engineer - Prediction and Planning ML
$151k+/yrOn-site3+ YOEML Engineering

Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.

Fab2

Fab2

Austin, TX
Software Engineer, AI Platform
$140k+/yrOn-siteML Engineering

Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.