Skip to content
RunwayRunway

Member of Technical Staff, Research Engineer (Datasets)

Research Engineer owning datasets for training world simulation AI models, designing multimodal datasets, running experiments, and building data pipelines to enhance model capabilities across tasks like robotics and creative tools. Requires 4+ years in ML with experience in generative models and frameworks like PyTorch or JAX.

About the job

What you'll do

  • Design multimodal, multitask datasets that teach world models new capabilities — deciding what data to collect, generate, or curate and measuring its effect on model behavior
  • Run controlled training experiments to understand how data composition drives model performance across tasks and domains
  • Build and operate large-scale pipelines for synthetic data generation, filtering, and quality control
  • Define evaluations and benchmarks that measure whether our models are actually improving at the things that matter
  • Partner with product and creative teams to translate target behaviors and capabilities into concrete data strategies

What you'll need

  • 4+ years of experience in machine learning, bonus points for data-centric approaches
  • Experience with large multimodal datasets and generative models (video, image, or multimodal)
  • Deep intuition for how data composition and quality translate to model capabilities
  • Comfort working across the full research stack: data analysis, dataset creation, model training, evaluation, and back again
  • Proficiency with at least one ML framework (e.g. PyTorch, JAX) and distributed compute tools (e.g. Ray, Kubernetes)
  • Excitement about building AI that simulates the world

Skills

PyTorch, JAX, Ray, Kubernetes, Machine Learning, Generative Models, Multimodal Datasets, Synthetic Data Generation, Distributed Compute, Data Pipelines

Fluidstack

Fluidstack

Austin, TX
Data Engineer
$269k+/yrOn-site5+ YOEData Engineering

Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Data Operations
$250k+/yrHybridData Engineering

Own end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.

OpenAI

OpenAI

Mountain View, CA
Data Engineer, Monetization Data Platform
$230k+/yrOn-siteData Engineering

Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.

The Voleon Group

The Voleon Group

New York, NY
Software Engineer, Strategy Research Analytics
$230k+/yrRemote3+ YOEData Engineering

Build and evolve reliable analytics infrastructure, pipelines, schemas, and foundational datasets supporting quantitative research across strategies. The role requires strong Python and SQL skills, distributed data-platform experience, and ownership of observability, performance, and reproducibility.

Anthropic

Anthropic

San Francisco, CA
Data Engineer, GTM
$320k+/yrHybrid5+ YOEData Engineering

Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.