Research Scientist / Engineer — Multimodal Agent
Builds and trains large-scale multimodal agentic models involving reasoning, planning, coding, and tool calling. Requires strong ML foundations, PyTorch expertise, and experience with distributed training on massive datasets.
About the job
What You'll Do
Modeling
- Architect large-scale multimodal agentic models that use reasoning, planning, coding, and tool calling to achieve complex, multi-step multimodal work.
Data
- Hillclimbing existing tasks and formulating new tasks through data.
- Design, implement, and run robust data pipelines for constructing, enriching, and filtering massive pixel datasets.
Systems
- Train large-scale multimodal models on massive datasets and GPU clusters.
Evaluation
- Define and build novel evaluation frameworks to measure multimodal agents.
Who You Are
- Strong foundation in machine learning, foundation models and agentic systems.
- Deep understanding of agentic systems and approaches in LLM/VLM reasoning, coding models, LLM/VLM tool calling.
- Hands-on experience with PyTorch and large-scale training (distributed, mixed precision, large datasets).
What Sets You Apart (Bonus Points)
- Experience in the following around data, modeling, or evaluation: State-of-the-art foundation models in reasoning, State-of-the-art foundation models in coding, State-of-the-art foundation models in tool calling, State-of-the-art multimodal agents.
Compensation
- The base pay range for this role is $250,000 – $450,000 per year.
Skills
PyTorch, Machine Learning, Foundation Models, LLMs, Vlm, Distributed Training, Mixed Precision Training, Multimodal Models, Agentic Systems, Reasoning Models, Coding Models, Tool Calling, Data Pipelines, Gpu Clusters
Similar jobs
ML Engineering jobsBuild and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.
Build and deploy algorithmic systems for high-impact healthcare problems, choosing among machine learning, optimization, heuristics, and hybrid approaches. The role requires 4+ years of relevant industry experience, strong applied problem-solving and evaluation skills, and fluency in modern ML tooling.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build and optimize the production LLM inference runtime for frontier models on OpenAI’s custom silicon. The role spans scheduling, distributed execution, memory and KV-cache management, performance tooling, and hardware-software co-design.
Build and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.