Head of Research
Leads the research agenda and hands-on development of replayable enterprise environments, agent evaluations, and post-training systems. The role requires deep AI research experience, a PhD or equivalent track record, and the ability to translate open-ended questions into production systems.
About the job
Responsibilities
- Own the research agenda for building a replayable environment engine over real enterprise history.
- Identify high-leverage technical questions and design experiments to answer them.
- Build systems that convert recorded enterprise data and task definitions into runnable environments.
- Design graders that turn ambiguous business objectives into verifiable rewards.
- Mine useful tasks, trajectories, and evaluation cases from historical workflows.
- Create representative, reproducible, and overfitting-resistant evaluation sets.
- Determine effective combinations of models, tools, context, and policies while reducing inference cost.
- Advance post-training methods for agents operating over long horizons, incomplete information, and large tool spaces.
- Build replay and observability systems that make agent behavior explainable and measurable.
- Scale to thousands of concurrent training and evaluation runs.
- Help establish research practices, evaluate technical progress, select research bets, and recruit and develop the research team.
- Collaborate directly with the CTO and deploy research into enterprise workflows.
Requirements
- PhD in machine learning, computer science, mathematics, or equivalent significant research experience.
- Deep experience in reinforcement learning, LLM post-training, evaluations, agent environments, or closely related areas.
- Experience taking ambitious, open-ended research problems from hypothesis through experimentation to working systems.
- Ability to turn ambiguous business objectives into reliably evaluable tasks and signals.
- Ability to move between research and production implementation.
Nice-to-haves
- Experience at a leading foundation model lab, top AI research organization, or high-performing AI startup.
Compensation and Benefits
- Salary: $250,000–$400,000 per year.
- Significant equity and ownership.
- Equinox membership.
- Free meals, coffee, and snacks.
- Health insurance.
- Unlimited PTO.
Skills
Reinforcement Learning, Llm Post-Training, Evaluation Systems, Agent Environments, Machine Learning, Python, Context Engineering, Agent Engineering, Observability, Open-Weight Models
Similar jobs
AI Research jobsLeads Deepgram’s end-to-end TTS research program, setting technical direction, training and evaluating large-scale speech-generation models, and turning breakthroughs into production systems. The role combines hands-on technical leadership with building and developing a high-performing research organization.
Sets company-wide architecture and strategy for data and applied AI, connecting governed data foundations to production intelligence and measurable business outcomes. The role requires 14+ years of experience, strong production engineering judgment, executive partnership, and hands-on delivery.
Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.
Leads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.
Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.