Member of Technical Staff, Vision / Language
Research Engineer/Scientist building vision-language pipelines and data systems for robot learning. Focus on turning egocentric/teleoperation video into high-signal training data for VLAs and world models.
About the job
What You'll Do
- Design and implement vision-language pipelines for egocentric and teleoperation video: structured captioning, temporal grounding, action-conditioned scene understanding, and semantic annotation at scale
- Develop and evaluate representations that bridge visual perception, language, and low-level robot action — spanning VLAs, video prediction, and world models
- Build and improve data curation systems that assess quality, diversity, and coverage of large-scale robot demonstration datasets
- Work hands-on with bimanual and high-DoF manipulation data, including real teleoperation footage and sim-generated rollouts
- Collaborate directly with partner labs to define data requirements and close the loop between data quality and downstream policy performance
- Stay current on the research frontier (VLAs, video foundation models, flow matching, DiT architectures, egocentric pretraining) and translate insights into production systems
Requirements
- MS or PhD in Computer Science, Robotics, Machine Learning, or a related field from a top-tier program
- 3–7 years of research or applied research experience (industry or academic) in one or more of: vision-language models, video understanding, robot learning, or generative modeling
- Deep fluency in PyTorch; working knowledge of large-scale training infrastructure (distributed training, mixed precision, large batch workflows)
- Published work or demonstrable impact in VLMs/VLAs, video representation learning, imitation learning, or a closely related area
- Strong engineering fundamentals — you can design clean systems, not just run experiments
Benefits
- Competitive compensation and equity
- Comprehensive health and wellness benefits
- Flexible work arrangements
- Collaborative and fast-paced work environment
- Opportunity to shape the future of robotics and AI alongside an ambitious, values-driven team
Skills
PyTorch, Vision-Language Models, Video Understanding, Robot Learning, Generative Modeling, Distributed Training, Mixed Precision Training, Vlas, Video Prediction, World Models
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.