Research Scientist, Frontier Risk Evaluations
Designs evaluation measures, harnesses, and datasets to assess risks from frontier AI systems, including dangerous capabilities testing. Collaborates with agencies, publishes methodologies for policymakers; requires 3+ years ML experience and publications in generative AI.
About the job
Responsibilities
- Design and build harnesses to test AI models and systems (including agents) for dangerous capabilities such as security vulnerability exploitation, CBRN uplift, and other high-risk activities.
- Work with government agencies or other labs to collectively scope and design evaluations to measure and mitigate risks posed by advanced AI systems.
- Publish evaluation methodologies and write technical reports for policymakers.
Requirements
- Commitment to promoting safe, secure, and trustworthy AI deployments.
- Practical experience conducting technical research collaboratively, including building and instrumenting ML pipelines, writing evaluation harnesses, and prototyping ideas from research literature.
- Track record of published research in machine learning, particularly generative AI.
- At least three years of experience addressing sophisticated ML problems in research or product development.
- Strong written and verbal communication skills for cross-functional teams.
Nice to Have
- Experience crafting evaluations and benchmarks, or background in data science roles related to LLM technologies.
- Experience with red-teaming or adversarial testing of AI systems.
- Familiarity with AI safety policy frameworks (e.g., NIST AI RMF, EU AI Act, Korea AI Basic Act).
Skills
Machine Learning, Generative AI, LLMs, Ml Pipelines, Evaluation Harnesses, Red-Teaming, Adversarial Testing, Ai Safety, Benchmarks, Prototyping
Similar jobs
AI Research jobsConduct rigorous people research and applied data science to evaluate talent programs, organizational health, and employee experiences. The role requires advanced expertise in research design, experimentation, measurement, causal inference, statistical modeling, and responsible handling of sensitive employee data.
Leads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.
Research Scientist defining and executing research on reliable long-horizon agents in enterprise environments. The role focuses on post-training and reinforcement learning, agent memory, evaluation, verification, and structured representations, combining hands-on experimentation with product delivery and publication.
Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.
Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.