Research Engineer
Leads pre-training and post-training of action-conditioned world models and VLA models for physical AI applications. Requires PyTorch expertise, distributed training, and ML fundamentals; robotics background preferred.
About the job
Responsibilities
- Design, implement, and run pre-training and post-training pipelines for action-conditioned world models and vision-language-action (VLA) models
- Develop and refine training methodologies, including fine-tuning, reinforcement learning, and large-scale multimodal learning
- Design and generate training and evaluation datasets from simulation, including environment setup, domain randomization, and sim-to-real transfer strategies
- Build distributed training infrastructure using PyTorch, FSDP, and DeepSpeed
- Work with multimodal data pipelines involving video, sensory inputs, and action sequences
- Evaluate model performance using both benchmark datasets and real-world deployment metrics
- Contributions to research publications a plus
- Collaborate with industrial partners to adapt generative models for real-world physical AI applications
Qualifications
- Experience with pre-training or post-training on large generative models (video, multimodal, or action-conditioned)
- Hands-on proficiency with PyTorch and distributed training frameworks (FSDP, DeepSpeed)
- Strong fundamentals in machine learning, optimization, and large-scale data processing
- Familiarity with VLMs, VLAs, or world models
- Background in robotics, embodied AI, or sim-to-real transfer is a plus
- Experience with video understanding or temporal reasoning is a plus
- BS/MS/PhD in Computer Science, Machine Learning, Robotics, or a related field
Benefits
- Competitive compensation and equity
- 401k (no match)
- Healthcare (Silver PPO Medical, Vision, Dental)
- Lunch and snacks at the office
Skills
PyTorch, Fsdp, Deepspeed, Reinforcement Learning, Multimodal Learning, Vision-Language-Action Models, World Models, Sim-To-Real Transfer, Video Understanding, Vlms
Similar jobs
ML Engineering jobsBuild and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.
Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.
Build evaluation methods, RL environments, agent tooling, and scalable infrastructure that make subjective qualities such as design and taste measurable for frontier AI models. The role requires experience with evaluations, RL environments, ML or post-training, plus strong backend engineering skills.
Build and scale generative video and multimodal models, optimizing training and inference for efficiency, throughput, and ultra-low latency. The role requires deep learning systems expertise, strong PyTorch/CUDA experience, and the ability to move research models into production.
Build and ship production AI agents and the platform infrastructure that makes them reliable, steerable, and measurable. The role requires strong backend fundamentals, production LLM or agent experience, and expertise in evaluations, retrieval, orchestration, or tool-use design.