Skip to content
AmbralAmbral

Head of Research

Leads the research agenda and hands-on development of replayable enterprise environments, agent evaluations, and post-training systems. The role requires deep AI research experience, a PhD or equivalent track record, and the ability to translate open-ended questions into production systems.

About the job

Responsibilities

  • Own the research agenda for building a replayable environment engine over real enterprise history.
  • Identify high-leverage technical questions and design experiments to answer them.
  • Build systems that convert recorded enterprise data and task definitions into runnable environments.
  • Design graders that turn ambiguous business objectives into verifiable rewards.
  • Mine useful tasks, trajectories, and evaluation cases from historical workflows.
  • Create representative, reproducible, and overfitting-resistant evaluation sets.
  • Determine effective combinations of models, tools, context, and policies while reducing inference cost.
  • Advance post-training methods for agents operating over long horizons, incomplete information, and large tool spaces.
  • Build replay and observability systems that make agent behavior explainable and measurable.
  • Scale to thousands of concurrent training and evaluation runs.
  • Help establish research practices, evaluate technical progress, select research bets, and recruit and develop the research team.
  • Collaborate directly with the CTO and deploy research into enterprise workflows.

Requirements

  • PhD in machine learning, computer science, mathematics, or equivalent significant research experience.
  • Deep experience in reinforcement learning, LLM post-training, evaluations, agent environments, or closely related areas.
  • Experience taking ambitious, open-ended research problems from hypothesis through experimentation to working systems.
  • Ability to turn ambiguous business objectives into reliably evaluable tasks and signals.
  • Ability to move between research and production implementation.

Nice-to-haves

  • Experience at a leading foundation model lab, top AI research organization, or high-performing AI startup.

Compensation and Benefits

  • Salary: $250,000–$400,000 per year.
  • Significant equity and ownership.
  • Equinox membership.
  • Free meals, coffee, and snacks.
  • Health insurance.
  • Unlimited PTO.

Skills

Reinforcement Learning, Llm Post-Training, Evaluation Systems, Agent Environments, Machine Learning, Python, Context Engineering, Agent Engineering, Observability, Open-Weight Models

Deepgram

Deepgram

San Francisco, CA
Director of Research, Text to Speech
$213k+/yrRemote8+ YOEAI Research

Leads Deepgram’s end-to-end TTS research program, setting technical direction, training and evaluating large-scale speech-generation models, and turning breakthroughs into production systems. The role combines hands-on technical leadership with building and developing a high-performing research organization.

Order.co

Order.co

Boston, MA
Principal Applied AI Architect
No salary listedRemote14+ YOEAI Research

Sets company-wide architecture and strategy for data and applied AI, connecting governed data foundations to production intelligence and measurable business outcomes. The role requires 14+ years of experience, strong production engineering judgment, executive partnership, and hands-on delivery.

Shield AI

Shield AI

Washington, DC
Senior Staff Engineer, Autonomy Capabilities – Maritime
$221k+/yrOn-site10+ YOEAI Research

Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.

Shield AI

Shield AI

San Mateo, CA

Senior Staff Software Engineer, Autonomy Capabilities
$281k+/yrOn-site10+ YOEAI Research

Leads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.

Upstart

Upstart

United States

Staff Machine Learning Model Risk Specialist
$140k+/yrRemote7+ YOEAI Research

Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.