Senior Machine Learning Engineer, Personalization, Magenta
Build and improve core agentic capabilities (memory, context, multi-step tool use) and evaluation frameworks (LLM-as-judge) for Spotify's Talk to Spotify conversational AI. Requires 5+ years production ML experience, rigorous evaluation skills, and comfort with rapid iteration on real user data.
About the job
What You'll Do
- Build and improve the core agentic capabilities that power the agent behind Talk to Spotify (memory, context management, multi-step tool use)
- Design and calibrate evaluation frameworks (including LLM-as-judge) that accelerate our confidence in the agent's behavior, and increase our offline-to-online success
- Work in a very dynamic space: prototype, dogfood, ship, learn, and refine in tight loops with real users as our understanding of the problem and users' expectations of agentic products and Spotify evolve
Who You Are
- Excited by agentic experiences — building agents, evaluating agents, and the hard problems in between (context handling, multi-step reasoning, ambiguity at scale)
- Like getting your hands dirty: shipping quickly, testing ideas against real usage, and learning from the wild rather than over-indexing on offline evaluation
- Have 5+ years of production ML experience deploying highly impactful products, or equivalent experience in other roles with a deep ML background
- Know how to evaluate ML systems rigorously — designing metrics, building eval pipelines, judge alignment, and can develop intuition through dogfooding and looking at user behavior
- Comfortable debugging the messy interactions between models, tools, and system constraints like latency
Skills
Machine Learning, LLMs, Agentic Systems, Evaluation Frameworks, Llm-As-Judge, Context Management, Multi-Step Reasoning, Multi-Step Tool Use, Metrics Design, Eval Pipelines, Python
Similar jobs
ML Engineering jobsBuild and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.
Build and deploy generative AI and LLM-powered agentic applications at Front to automate customer support inquiries, enhance product capabilities, and drive operational insights. Requires 5+ years software engineering experience with strong production AI/ML focus, agentic/RAG expertise, and proficiency in Node.js, TS, and Python.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.
Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.