Research Engineer, Frontier Evals & Environments
Builds ambitious RL environments and evaluation systems to measure and steer frontier AI models toward safe AGI. Requires strong ML research engineering, statistical skills, and red-teaming mindset for end-to-end project ownership in fast-paced setting.
About the job
Responsibilities
- Create ambitious RL environments to push our models to their limits
- Work on measuring frontier model capabilities, skills, and behaviors
- Develop new methodologies for automatically exploring the behavior of these models
- Help steer training for our largest training runs, and see the future first
- Design scalable systems and processes to support continuous evaluation
- Build self-improvement loops to automate model understanding
Requirements
- Passionate and knowledgeable about AGI/ASI measurement
- Strong engineering and statistical analysis skills
- Able to think outside the box and have a robust “red-teaming mindset”
- Experienced in ML research engineering, stochastic systems, observability and monitoring, LLM-enabled applications, and/or another technical domain applicable to AI evaluations
- Able to operate effectively in a dynamic and extremely fast-paced research environment as well as scope and deliver projects end-to-end
Nice-to-haves
- First-hand experience in red-teaming systems—be it computer systems or otherwise
- An ability to work cross-functionally
- Excellent communication skills
Skills
Reinforcement Learning, Machine Learning, LLMs, Statistical Analysis, Red-Teaming, Observability, Monitoring, Stochastic Systems, Rl Environments, Model Evaluation
Similar jobs
AI Research jobsLeads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.
Research Scientist defining and executing research on reliable long-horizon agents in enterprise environments. The role focuses on post-training and reinforcement learning, agent memory, evaluation, verification, and structured representations, combining hands-on experimentation with product delivery and publication.
Conduct rigorous people research and applied data science to evaluate talent programs, organizational health, and employee experiences. The role requires advanced expertise in research design, experimentation, measurement, causal inference, statistical modeling, and responsible handling of sensitive employee data.
Build and ship agentic AI product experiences, internal automation, and customer-facing features across the stack. The role requires 5+ years of software engineering experience, hands-on experience with AI or LLM-powered products, Python proficiency, and strong autonomy.
Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.