Agent Post-Training, Frontier Evals and Environments Research
Researcher building frontier RL environments, evaluations, and training signals to steer OpenAI's largest agent training runs and measure model capabilities.
About the job
Responsibilities
- Create ambitious RL environments to push frontier models to their limits and measure model capabilities, skills, and behaviors
- Develop new methodologies for automatically exploring model behavior
- Dive deep into the science of measurement, including scalability, reliability, and variance of evaluation methodology
- Help steer training for the largest training runs
- Design scalable systems and processes to support continuous evaluation
- Build self-improvement loops to automate model understanding
Requirements
- Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field
- Hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems
- Ability to move from a vague behavioral problem to a concrete experiment: define the hypothesis, build the pipeline, run the model, analyze the result, and decide next steps
- Comfortable working across research, product, infrastructure, data, evals, and safety boundaries
Nice-to-Haves
- Excitement for open-ended problems where the path is unclear and the signal is noisy
- Care about product impact and model behavior beyond benchmark movement
- Opinions about what makes an agent useful, reliable, honest, tasteful, and easy to work with
- Willingness to build load-bearing systems and processes even when the work is not glamorous
Skills
Machine Learning, Software Engineering, Statistics, LLMs, Reinforcement Learning, RLHF, Rlaif, Post-Training, Evaluations, Graders, Synthetic Data, Model Training, Coding Agents, Tool-Using Agents, Production Ml Systems
Similar jobs
AI Research jobsApplied AI Research Engineer who tests model capabilities, builds demos and evaluations, supports strategic customer implementations, and translates field insights into product and research direction. Requires 6+ years of technical experience, programming proficiency, LLM development experience, and strong communication skills.
Leads the research agenda for humanoid robotics, developing foundation-model and reinforcement-learning methods for dexterous manipulation and deploying them on real robotic systems. Requires a PhD, strong robotics research publications, and senior-level technical leadership.
Research and build safety models, evaluations, and runtime safeguards for conversational AI agents, addressing prompt injection, unsafe tool use, privacy, and policy risks. Requires 4+ years in AI/ML engineering, research, or safety plus experience deploying and evaluating language models or agentic systems.
Build and expand customer-facing agentic AI products, MCP integrations, and automated reconciliation workflows for private fund management. The role requires senior-level software engineering, strong systems thinking, product judgment, and hands-on experience building and evaluating AI systems.
Build Vanta’s organizational intelligence layer by shipping prototypes, internal tools, and AI agent workflows that make cross-source data useful to EPD, GTM, and other teams. The role requires recent hands-on LLM product work, independent problem scoping, and strong judgment around AI quality, reliability, cost, and latency.