Applied AI Researcher, Benchmarking
Designs and constructs AI benchmarks and evaluation frameworks to measure reasoning, reliability, and real-world impact of intelligent systems. Requires experience with model evaluations, statistical rigor, building with AI models, and strong programming for prototypes.
About the job
Key Responsibilities
- Design evaluation frameworks that capture reasoning depth, interaction quality, reliability, and operational impact.
- Construct benchmarks that reflect real-world complexity to judge new architectures, techniques, and releases.
- Explore new paradigms for evaluating intelligent systems: adversarial robustness testing, longitudinal performance tracking, and human-in-the-loop assessment.
- Investigate how metrics shape model behavior and establish rigorous methodologies for quantifying emergent capability.
Who You Are (Requirements)
- Experience designing and running evaluations: built or maintained benchmarks, test suites, or experimental frameworks.
- Statistical and analytical rigor: design fair, reproducible experiments and extract signal from noisy results.
- Experience building with models (compound AI systems, agentic collaboration, ensembling, ReAct, graph-of-thoughts, etc.).
- Proven track record of research results (publications, public work).
- Uses AI every day (ChatGPT, Cursor, Perplexity).
- Strong programming and data analysis skills for prototypes and experiments.
- Biases towards showing vs telling.
Compensation & Benefits
- Base salary: $150K – $250K (depending on experience, location, level).
- Meaningful equity.
- 100% covered medical, dental, vision for employees/dependents.
- 401(k), commuter benefits, in-office lunch.
- Access to state-of-the-art models and AI tools.
Skills
Ai Benchmarks, Evaluation Frameworks, Llm Evaluation, React, Graph-Of-Thoughts, Ensembling, Python, Data Analysis, Compound Ai Systems, Agentic Systems
Similar jobs
AI Research jobsConduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.
Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.
Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.
Research Engineer focused on designing benchmarks, evaluation systems, rubrics, and failure-analysis workflows for frontier language models. The role requires strong applied AI research and coding experience, with expertise in model evaluation, data quality, and backend systems.
Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.