Member of Technical Staff, Research Engineer (Datasets)
Research Engineer owning datasets for training world simulation AI models, designing multimodal datasets, running experiments, and building data pipelines to enhance model capabilities across tasks like robotics and creative tools. Requires 4+ years in ML with experience in generative models and frameworks like PyTorch or JAX.
About the job
What you'll do
- Design multimodal, multitask datasets that teach world models new capabilities — deciding what data to collect, generate, or curate and measuring its effect on model behavior
- Run controlled training experiments to understand how data composition drives model performance across tasks and domains
- Build and operate large-scale pipelines for synthetic data generation, filtering, and quality control
- Define evaluations and benchmarks that measure whether our models are actually improving at the things that matter
- Partner with product and creative teams to translate target behaviors and capabilities into concrete data strategies
What you'll need
- 4+ years of experience in machine learning, bonus points for data-centric approaches
- Experience with large multimodal datasets and generative models (video, image, or multimodal)
- Deep intuition for how data composition and quality translate to model capabilities
- Comfort working across the full research stack: data analysis, dataset creation, model training, evaluation, and back again
- Proficiency with at least one ML framework (e.g. PyTorch, JAX) and distributed compute tools (e.g. Ray, Kubernetes)
- Excitement about building AI that simulates the world
Skills
PyTorch, JAX, Ray, Kubernetes, Machine Learning, Generative Models, Multimodal Datasets, Synthetic Data Generation, Distributed Compute, Data Pipelines
Similar jobs
Data Engineering jobsBuild and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.
Own end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.
Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.
Build and evolve reliable analytics infrastructure, pipelines, schemas, and foundational datasets supporting quantitative research across strategies. The role requires strong Python and SQL skills, distributed data-platform experience, and ownership of observability, performance, and reproducibility.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.