Skip to content
OpenAIOpenAI

Research Engineer, Frontier Evals & Environments

Builds ambitious RL environments and evaluation systems to measure and steer frontier AI models toward safe AGI. Requires strong ML research engineering, statistical skills, and red-teaming mindset for end-to-end project ownership in fast-paced setting.

About the job

Responsibilities

  • Create ambitious RL environments to push our models to their limits
  • Work on measuring frontier model capabilities, skills, and behaviors
  • Develop new methodologies for automatically exploring the behavior of these models
  • Help steer training for our largest training runs, and see the future first
  • Design scalable systems and processes to support continuous evaluation
  • Build self-improvement loops to automate model understanding

Requirements

  • Passionate and knowledgeable about AGI/ASI measurement
  • Strong engineering and statistical analysis skills
  • Able to think outside the box and have a robust “red-teaming mindset”
  • Experienced in ML research engineering, stochastic systems, observability and monitoring, LLM-enabled applications, and/or another technical domain applicable to AI evaluations
  • Able to operate effectively in a dynamic and extremely fast-paced research environment as well as scope and deliver projects end-to-end

Nice-to-haves

  • First-hand experience in red-teaming systems—be it computer systems or otherwise
  • An ability to work cross-functionally
  • Excellent communication skills

Skills

Reinforcement Learning, Machine Learning, LLMs, Statistical Analysis, Red-Teaming, Observability, Monitoring, Stochastic Systems, Rl Environments, Model Evaluation

Mercor

Mercor

San Francisco, CA

Research Scientist, APEX Benchmarks
$200k+/yrOn-siteAI Research

Leads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.

Tessera Labs

Tessera Labs

San Jose, CA

Research Scientist
$200k+/yrOn-siteAI Research

Research Scientist defining and executing research on reliable long-horizon agents in enterprise environments. The role focuses on post-training and reinforcement learning, agent memory, evaluation, verification, and structured representations, combining hands-on experimentation with product delivery and publication.

OpenAI

OpenAI

San Francisco, CA

People Research Scientist
$198k+/yrOn-siteAI Research

Conduct rigorous people research and applied data science to evaluate talent programs, organizational health, and employee experiences. The role requires advanced expertise in research design, experimentation, measurement, causal inference, statistical modeling, and responsible handling of sensitive employee data.

Baseten

Baseten

San Francisco, CA

AI Engineer
$220k+/yrHybrid5+ YOEAI Research

Build and ship agentic AI product experiences, internal automation, and customer-facing features across the stack. The role requires 5+ years of software engineering experience, hands-on experience with AI or LLM-powered products, Python proficiency, and strong autonomy.

Earnin

Earnin

Mountain View, CA

Software Engineer (Gen AI)
$181k+/yrHybrid3+ YOEAI Research

Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.