Research Scientist, Agent Robustness
Research Scientist focuses on agent robustness, developing tests, exploits, and mitigations for safe AI agents. Requires 3+ years ML experience, RL techniques like RLHF/DPO, and published research in generative AI.
About the job
Responsibilities
- Research the science of AI agent capabilities with a focus on safety, risk factors, and benchmarking methodologies.
- Design and build harnesses to test AI agents’ tendency to take harmful actions when pressured or tricked.
- Design and build exploits and mitigations for failure modes arising from agent affordances like coding, web browsing, and computer use.
- Characterize and design mitigations for failure modes or risks in systems with multiple interacting AI agents.
Requirements
- Commitment to promoting safe, secure, and trustworthy AI deployments.
- Practical experience conducting technical research collaboratively, including building agent scaffolding, designing evaluation harnesses, and prototyping research ideas.
- Experience with post-training and RL techniques such as RLHF, DPO, GRPO.
- Track record of published research in machine learning, particularly generative AI.
- At least 3 years of experience addressing sophisticated ML problems.
- Strong written and verbal communication skills.
Nice to Have
- Hands-on experience with agent evaluation frameworks such as SWE-bench, WebArena, OSWorld, Inspect.
- Experience with red-teaming, prompt injection, or adversarial testing of AI systems.
Skills
RLHF, Dpo, Grpo, Swe-Bench, Webarena, Osworld, Inspect, Red-Teaming, Prompt Injection, Adversarial Testing, Generative AI, Machine Learning, Agent Evaluation, Rl Techniques
Similar jobs
AI Research jobsConduct rigorous people research and applied data science to evaluate talent programs, organizational health, and employee experiences. The role requires advanced expertise in research design, experimentation, measurement, causal inference, statistical modeling, and responsible handling of sensitive employee data.
Leads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.
Research Scientist defining and executing research on reliable long-horizon agents in enterprise environments. The role focuses on post-training and reinforcement learning, agent memory, evaluation, verification, and structured representations, combining hands-on experimentation with product delivery and publication.
Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.
Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.