AI Research Scientist, New Grad – Agents & Reinforcement Learning
Conduct research on autonomous AI agents and reinforcement learning to build self-improving systems that reason, code, and learn at scale within the Snowflake Data Cloud. Requires a PhD (or equivalent) and strong expertise in RL and agentic AI.
About the job
Responsibilities
- Design and develop agentic frameworks powered by recursive self-improvement loops, enabling AI systems that iteratively refine their own capabilities and strategies
- Build and evaluate auto research agents — systems capable of autonomously formulating hypotheses, executing experiments, and synthesizing findings
- Develop coding agents that understand, generate, and debug code across complex, multi-step programming tasks
- Conduct research in reinforcement learning with a focus on RLHF, DPO, and PPO as mechanisms for aligning and improving agentic behaviors
- Contribute to multi-agent systems where specialized agents collaborate, negotiate, and self-organize to solve enterprise-scale problems
- Develop and curate training data pipelines — both synthetic and human-annotated — to support novel agentic and RL research domains
- Publish research findings at top-tier venues such as NeurIPS, ICML, ICLR, and ACL
Requirements
- PhD in Computer Science, Machine Learning, Artificial Intelligence, or a closely related field (completing or recently completed; or equivalent research experience)
- Foundational expertise in reinforcement learning algorithms, including RLHF, DPO, PPO, or multi-agent systems
- Research experience in LLM post-training, fine-tuning, or reasoning model development
- Demonstrated ability to implement and experiment with agentic architectures — including tool-use, planning, and self-correction loops
- Proficiency in Python and at least one deep learning framework (PyTorch or JAX strongly preferred)
- Strong mathematical and analytical foundation — comfortable working at the intersection of theory and empirical research
- At least one first-author or co-authored publication or preprint in a relevant AI/ML area
Nice-to-Haves
- Hands-on experience building or evaluating coding agents or auto research agents
- Familiarity with recursive self-improvement frameworks or automated AI scientist paradigms
- Experience with large-scale distributed training or efficient training paradigms
- Background in mathematical reasoning, structured decision-making, or program synthesis
- Exposure to domain-specific AI applications in healthcare, finance, or enterprise workflows
Skills
Python, PyTorch, JAX, Reinforcement Learning, RLHF, Dpo, Ppo, Multi-Agent Systems, Llm Post-Training, Agentic Architectures
Similar jobs
AI Research jobsConduct foundational research on LLMs and multimodal systems, designing architectures and training methods and helping move prototypes into production. The role targets PhD researchers graduating by December 2026 with strong machine-learning research and programming experience.
Research Engineer developing and deploying machine-learning algorithms for autonomous driving and robotics systems. The role targets recent MS or PhD graduates with experience in areas such as foundation models, diffusion policies, reinforcement learning, computer vision, and robotics.
Build and evolve the agent harness powering Perplexity’s flagship answer experience, improving orchestration, context management, performance, reliability, observability, and evaluation. The role requires strong software engineering skills, Python proficiency, and experience shipping large-scale AI systems.
Leads experiments investigating working-memory circuits in behaving mice through multiregional Neuropixels recordings, optical perturbation, and large-scale neural-data analysis. Requires PhD-level neuroscience training, mouse survival surgery, awake-rodent electrophysiology, and Python or MATLAB expertise.
Conduct research and develop foundation models for robotic manipulation and high-precision manufacturing, taking projects from data curation through deployment on industrial robots. The role requires current PhD study, strong Python and deep learning expertise, robotics simulation experience, and research in foundation models.