Machine Learning Researcher, Multimodal LLMs
Develops next-generation multimodal LLMs integrating speech, text, tools, and real-time reasoning for conversational AI agents. Requires strong background in LLMs, multimodal models, fast experimentation, and production deployment experience.
About the job
Responsibilities
- Contribute to the development of next-generation multimodal LLM stack, combining speech, text, tools, and real-time reasoning into a single unified system.
- Build industry-leading conversational AI models that power Bland's agent, taking them from idea to production.
- Define how agents listen, think, and act in real time, integrating streaming audio, tool execution, and dynamic context.
Requirements
- Strong LLM / Multimodal Background: Experience with LLMs, multimodal models, or speech-language systems. Deep understanding of prompting, fine-tuning, and alignment techniques. Familiarity with neural audio codecs and modern multimodal LLM techniques.
- Fast Experimental Loop: Ability to go from idea → dataset → experiment → conclusion in days. Design experiments that answer questions.
- Product Intuition: Strong sense for natural vs robotic interactions. Translate abstract modeling ideas into user-facing improvements.
- Builder Mentality: Take ownership from research through deployment. Thrive in ambiguous, fast-moving environments. Care about impact over elegance.
- Think in systems, obsess over latency, correctness, and real-world behavior. Comfortable discarding ideas quickly. Push toward simple abstractions.
Nice-to-Haves (Bonus Points)
- Experience with real-time voice systems or conversational AI.
- Background in tool-using agents or agent frameworks.
- Experience with multimodal datasets (audio + text + actions).
- Contributions to LLM or speech-related research or open source.
Compensation & Benefits
- Competitive salary: $180,000 – $260,000
- Meaningful equity
- Full healthcare, dental, vision
Skills
LLMs, Multimodal Models, Speech-Language Systems, Prompting, Fine-Tuning, Alignment Techniques, Neural Audio Codecs, Conversational AI, Real-Time Voice Systems, Tool-Using Agents
Similar jobs
AI Research jobsResearch Engineer focused on designing benchmarks, evaluation systems, rubrics, and failure-analysis workflows for frontier language models. The role requires strong applied AI research and coding experience, with expertise in model evaluation, data quality, and backend systems.
Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.
Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.
Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.
Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.