Applied Machine Learning Researcher
Conduct applied research and engineering to improve language-model behavior in real-time voice conversations. The role focuses on fine-tuning, rigorous evaluation, production failure analysis, data strategies, and safely deploying improvements.
About the job
Responsibilities
- Research and improve LLM behavior for real-time voice conversations.
- Design and run fine-tuning experiments across data, model, and evaluation strategies.
- Build evaluation frameworks for model quality, workflow-following, naturalness, reliability, task completion, and overall conversation quality.
- Analyze production conversations to identify failure modes and improvement opportunities.
- Develop data curation, labeling, and synthetic data strategies.
- Compare model architectures, training approaches, prompts, and datasets.
- Investigate regressions and clearly explain why model behavior improves or degrades.
- Work with engineering to deploy research improvements safely and efficiently.
- Help define model release criteria, evaluation gates, and quality benchmarks.
Requirements
- Strong experience with machine learning and modern language models.
- Hands-on experience fine-tuning, evaluating, or adapting language models.
- Strong Python skills and comfort working with messy, real-world datasets.
- Ability to design rigorous experiments and interpret results clearly.
- Experience building or improving evaluation systems for AI models.
- Strong analytical skills and ability to debug model behavior.
- Clear written and spoken communication with technical and non-technical teammates.
- Practical focus on production impact, latency, reliability, and customer outcomes.
Nice to Have
- PhD in machine learning, NLP, or a related field.
- Experience with conversational AI, voice AI, or customer-support automation.
- Experience with SFT, preference tuning, DPO, RFT, GRPO, RLHF, LoRA, or QLoRA.
- Familiarity with model serving, inference optimization, or vLLM.
- Experience with synthetic data generation and data quality pipelines.
- Experience in a startup or fast-moving product environment.
Skills
Machine Learning, LLMs, Python, Fine-Tuning, Model Evaluation, NLP, Sft, Dpo, RLHF, Lora, Qlora, vLLM, Synthetic Data, Model Serving, Inference Optimization
Similar jobs
AI Research jobsConduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.
Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.
Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.
Build Vanta’s organizational intelligence layer by shipping prototypes, internal tools, and AI agent workflows that make cross-source data useful to EPD, GTM, and other teams. The role requires recent hands-on LLM product work, independent problem scoping, and strong judgment around AI quality, reliability, cost, and latency.