Staff Machine Learning Research Scientist/ Engineer, Agents
Conducts cutting-edge research on data for state-of-the-art AI agents like browser and SWE agents, develops prototypes using LLMs and frameworks like PyTorch/JAX, and publishes in top ML venues. Requires 3+ years ML experience and strong cross-functional communication.
About the job
About This Role
This role is at the intersection of cutting-edge AI research and practical application, with a focus on studying the data types essential for building state-of-the-art agents, such as browser and SWE agents. The ideal candidate will explore the data landscape needed to advance intelligent, adaptable AI agents, guiding the data strategy at Scale to drive innovation. This position requires expertise in LLM agents and planning algorithms, creativity in addressing novel challenges related to data, interaction, and evaluation. You will contribute to impactful research publications on agents, collaborate with customer researchers, and work alongside the engineering team to translate these advancements into real-world, scalable solutions.
Ideally you’d have:
- Practical experience working with LLMs, with proficiency in frameworks like Pytorch, Jax, or Tensorflow. You should also be adept at interpreting research literature and quickly turning new ideas into prototypes.
- A track record of published research in top ML venues (e.g., ACL, EMNLP, NAACL, NeurIPS, ICML, ICLR, COLM, etc.)
- At least three years of experience addressing sophisticated ML problems, either in a research setting or product development.
- Strong written and verbal communication skills and the ability to operate cross-functionally.
Nice to have:
- Hands-on experience with open source LLM fine-tuning or involvement in bespoke LLM fine-tuning projects using Pytorch/Jax.
- Hands-on experience and publications in building applications and evaluations related to AI agents such as tool-use, text2SQL, browser agents, coding agents and GUI agents.
- Hands-on experience with agent frameworks such as OpenHands, Swarm, LangGraph, etc.
- Familiarity with agentic reasoning methods such as STaR and PLANSEARCH.
- Experience working with cloud technology stack (eg. AWS or GCP) and developing machine learning models in a cloud environment.
Skills
PyTorch, JAX, TensorFlow, LLMs, AI Agents, LangGraph, Openhands, Swarm, AWS, GCP
Similar jobs
AI Research jobsLeads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.
Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.
Research and evaluate frontier AI capabilities for cybersecurity, rapidly prototyping tools, designing rigorous benchmarks, and helping operationalize reliable capabilities into products. Requires deep security expertise, strong technical communication, and at least seven years of relevant experience.
Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.
Own the architecture, delivery, evaluation, and production operations of AI capabilities embedded in procurement and finance workflows. The role requires 10+ years in applied AI or machine learning, deep LLM and agent expertise, and experience delivering measurable production outcomes.