Student Worker - Machine Learning Engineer - Data Mining & VLM
Supports autonomous-vehicle road-rule compliance by building data miners and evaluating LLM/VLM triage pipelines across simulation and fleet data. Requires strong Python, PySpark, and SQL skills, multimodal model evaluation experience, and current enrollment in a relevant bachelor's or master's program.
About the job
Responsibilities
- Develop and iterate data miners that locate potential Rules of the Road violations while following SSO requirements.
- Develop and iterate VLM/LLM workflows and pipelines for classifying Rules of the Road events.
- Extend triagers from simulation data to fleet and service data.
- Curate golden datasets and evaluate triager precision and recall against human triage.
- Analyze error cases and translate findings into pipeline improvements.
Requirements
- Currently enrolled in a B.S. or M.S. program in a relevant discipline.
- Strong Python, PySpark, and SQL skills.
- Coursework in machine learning, computer vision, or natural language processing.
- Understanding of supervised learning, model evaluation, dataset bias, and label noise.
- Hands-on experience building production LLM/VLM applications, including agentic workflows, retrieval-augmented generation, tool calling, evaluation, or fine-tuning.
- Experience evaluating multimodal or vision-language models.
- Familiarity with Git, unit testing, debugging, and code reviews.
- Experience with prompt optimization, model fine-tuning, or human-in-the-loop machine learning systems.
- Available to commit to a minimum three-month assignment and at least 40 hours per week.
- Able to work on-site at one of the company's office locations.
Nice-to-Haves
- Strong Scala skills.
- Experience with Spark optimization and distributed data-processing systems.
- Experience with autonomous vehicles, robotics, mapping, or transportation-related datasets.
Compensation and Benefits
- Part-time student worker program.
- Minimum three-month assignment.
- Minimum commitment of 40 hours per week.
Skills
Python, Pyspark, SQL, Machine Learning, Computer Vision, Natural Language Processing, Llm Applications, Vlm Applications, RAG, Tool Calling, Model Evaluation, Fine-Tuning, Git, Scala, Spark
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.