Skip to content
PolymathPolymath

Member of Technical Staff - Research

Conducts applied research on long-horizon autonomous AI agents, focusing on evaluation, post-training, environment design, and benchmarks to improve frontier models. Builds simulations, runs experiments, ships production code, and publishes findings.

About the job

Responsibilities

  • Advance the frontier of autonomous agents through core research in long-horizon evaluation, agent post-training, and environment design.
  • Understand where current models fail and how to improve them.
  • Build benchmarks, create environments, write production code, and run rigorous experiments.
  • Develop advanced environment simulation engines for training & evaluating autonomous AI agents.
  • Investigate failure modes of frontier models.
  • Create rigorous benchmarks for complex, realistic tasks requiring long-horizon reasoning and tool use in dynamic environments.
  • Post-training agents in complex simulation environments.
  • Publish research.

Requirements

  • Strong engineering & research fundamentals and prolific user of AI tools.
  • Experience post-training frontier models.
  • Experience shipping reliable, production-quality code.
  • Track record of publications.

Perks

  • Comprehensive health, dental, and vision insurance.
  • 401(k).
  • Unlimited PTO.
  • Free meals with the team.
  • Wellness stipend & learning stipend.
  • Top of the line tech.
  • Frequent team activities and outings.

Skills

Reinforcement Learning, AI Agents, Simulation Environments, Post-Training, Frontier Models, Long-Horizon Reasoning, Benchmarks, Python, Machine Learning, Research Publication

Mercor

Mercor

San Francisco, CA

Research Scientist, APEX Benchmarks
$200k+/yrOn-siteAI Research

Leads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.

Tessera Labs

Tessera Labs

San Jose, CA

Research Scientist
$200k+/yrOn-siteAI Research

Research Scientist defining and executing research on reliable long-horizon agents in enterprise environments. The role focuses on post-training and reinforcement learning, agent memory, evaluation, verification, and structured representations, combining hands-on experimentation with product delivery and publication.

OpenAI

OpenAI

San Francisco, CA

People Research Scientist
$198k+/yrOn-siteAI Research

Conduct rigorous people research and applied data science to evaluate talent programs, organizational health, and employee experiences. The role requires advanced expertise in research design, experimentation, measurement, causal inference, statistical modeling, and responsible handling of sensitive employee data.

Earnin

Earnin

Mountain View, CA

Software Engineer (Gen AI)
$181k+/yrHybrid3+ YOEAI Research

Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.

Scale AI

Scale AI

San Francisco, CA
Machine Learning Research Scientist, Evaluations
$181k+/yrOn-siteAI Research

Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.