Senior Research Engineer
Research Engineer training and scaling flagship open multimodal and agentic models (Olmo, Molmo). Owns end-to-end ML infrastructure, model development, and open-source releases.
About the job
Key Responsibilities
- Building and optimizing infrastructure for LLM, multimodal, and agentic research — including training/inference pipelines, dataset curation, and large-scale preprocessing
- Designing, training, and evaluating multimodal models (vision + language) and agentic workflows, including tool use, planning, and long-horizon tasks
- Scoping and leading research projects, prioritizing experiments for highest impact
- Bringing strong software engineering practices to a research environment and bridging cutting-edge work to production-quality products
- Contributing to and supporting the open-source community through model releases, datasets, public APIs, and technical reports
Requirements
- 4+ years of ML infrastructure experience — data preprocessing, model training, evaluation, inference, and deployment
- Experience with end-to-end model development — dataset construction, training, fine-tuning, evaluation, profiling, and monitoring
- Familiarity with modern model architectures — including LLMs (MoEs, long-context models), vision-language models (e.g., Molmo, LLaVA), and experience training and evaluating both
- Agentic systems knowledge — tools, memory, and long-running workflows
- Strong software engineering fundamentals — performant, scalable systems and confident debugging
- Proficiency in Python and a major ML framework (PyTorch, JAX, or TensorFlow)
- Familiarity with cloud and containerization (e.g., GCP, AWS, Docker)
- Strong communication and collaboration skills
- BS or MSc in Computer Science, Statistics, Engineering, Applied Mathematics, or a related quantitative field (or equivalent experience)
- A minimum of 2 years of software development experience (or equivalent experience)
Nice-to-Haves
- PhD in ML or equivalent hands-on industry work with deep learning and/or foundation models
Skills
Python, PyTorch, JAX, TensorFlow, LLMs, Vision-Language Models, Moe, Docker, AWS, GCP
Similar jobs
ML Engineering jobsBuild marketplace search and ranking features while supporting MLOps infrastructure, model deployment, feature stores, and real-time data pipelines. The role requires 5+ years of software engineering or MLOps experience, backend or full-stack expertise, and familiarity with cloud and machine learning tooling.
Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.
Build and deploy production machine-learning models and data systems that classify and enrich Internet telemetry for internal platforms and customer-facing products. The role requires 5+ years of applied ML, data science, or software engineering experience, plus strong Python or Go skills.
Build and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
Leads hands-on development and deployment of production AI agents and workflows, sets technical standards, and mentors engineers. Requires 4+ years of software engineering experience, production LLM application experience, and strong Python or TypeScript skills.