Skip to content

Research Engineer - AI/RL Infrastructure

Designs, builds, and operates large-scale ML infrastructure for AI/RL research, including GPU cluster orchestration, data curation pipelines, and distributed training systems for autonomous driving and robotics.

About the job

Responsibilities

  • Design and build training and evaluation infrastructure to support AI research directions, orchestrating massive GPU clusters to process PBs of multimodal sensor data
  • Build robust benchmarking, continuous evaluation, and regression tracking systems to measure model performance across diverse, long-tail real-world driving distributions
  • Develop large-scale data sampling, dataset generation, and advanced data curation pipelines, leveraging state-of-the-art AI models to power a closed-loop data flywheel
  • Enable high-throughput distributed training across heterogeneous cloud environments, focusing on reliability, efficiency, and cost-aware scaling
  • Collaborate closely with AI research, autonomy, and platform teams to translate cutting-edge research into production-ready systems

Requirements

  • Experience building and operating production-grade software systems across the full machine learning lifecycle, including training, evaluation, data, and deployment
  • Opinions about building a company-wide platform for ML training, evaluation, and deployment
  • Experience with performance engineering and compute acceleration for large-scale ML training, including profiling, bottleneck analysis, and optimization
  • Strong systems-level debugging skills to diagnose and resolve issues in large-scale distributed training, spanning model code, data pipelines, runtimes, and cluster infrastructure
  • Deep familiarity with the open-source ML and systems ecosystem, with judgment on when to adopt open source versus build in-house
  • Technical experience in: PyTorch, CUDA, Ray, Flyte, Kubernetes

Nice to Have

  • Industry experience on relevant topics (self-driving application preferred)

Compensation

Base salary range: $126,000 - $423,000 USD annually, plus equity and benefits.

Skills

PyTorch, CUDA, Ray, Flyte, Kubernetes, Gpu Clusters, Distributed Training, ML Infrastructure, Data Pipelines, Performance Engineering

Build

Build

New York, NY
AI Engineer (Assistant)
$125k+/yrOn-siteML Engineering

Build and ship production agentic AI workflows for complex real estate and built-world processes. The role combines product engineering, applied AI, customer collaboration, workflow orchestration, evaluation, and reliable user-facing experiences.

Mintlify

Mintlify

San Francisco, CA

Applied AI Engineer
$130k+/yrOn-site4+ YOEML Engineering

Build and own customer-facing AI products from experimentation through production, including reliable agents, evaluation systems, APIs, interfaces, and infrastructure. Requires at least four years of software development experience and deep production experience with language-model systems.

PathAI

PathAI

Boston, MA
Machine Learning Engineer III
$131k+/yrOn-site5+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.

Build

Build

New York, NY

AI Engineer - Workflows
$120k+/yrOn-siteML Engineering

Build modular AI operations and evaluation systems that power complex real estate workflows. The role focuses on improving output quality, defining correctness with domain experts, and reducing human review while maintaining high standards.

Build

Build

New York, NY
AI Engineer - Assistant Experience
$120k+/yrOn-siteML Engineering

Build and operate Dougie, an agentic AI system that executes workflows, evaluates its own performance, retains institutional context, and improves in production. The role requires experience deploying unattended agentic systems and engineering reliable memory, retrieval, orchestration, and feedback loops.