Research Scientist, APEX Benchmarks
Leads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.
About the job
Responsibilities
- Design the next generation of APEX benchmarks, including task taxonomies, difficulty calibration, contamination controls, and statistical design.
- Create expert-built datasets and grading rubrics at scale while maintaining defensible quality standards.
- Establish rigorous measurement practices, including confidence intervals, inter-rater agreement, human-versus-model-as-judge calibration, held-out splits, and failure analysis.
- Collaborate with academic and industry partners to co-design and promote benchmark adoption.
- Publish research through papers, open datasets, blog posts, conference talks, and leaderboard releases.
- Represent the company's research with frontier labs, customers, partners, the press, and the broader research community.
- Translate benchmark findings into technical reports, customer narratives, and go-to-market materials.
- Partner with data operations, engineering, product, and strategy to move benchmarks from design into production and inform the company roadmap.
- Track LLM evaluation research and incorporate relevant methods into benchmark development.
Requirements
- Applied or academic research experience in LLM evaluation, benchmarking, NLP, or a related field.
- Track record of rigorous experimental design.
- Strong judgment about informative measurements of frontier-model behavior.
- Knowledge of sampling, variance, contamination, and grader reliability.
- Strong coding skills and ability to build evaluation harnesses, run experiments, and analyze results independently.
- Excellent written and verbal communication skills for technical and non-technical audiences.
- Comfort working in ambiguous, fast-moving, cross-functional environments.
- Interest in GTM strategy, startup dynamics, and the AI data business.
- Willingness to work onsite in the San Francisco office five days per week.
Nice to Have
- Ph.D. in machine learning, NLP, or a related field, or equivalent industry or frontier-lab research experience.
- Publications at top-tier venues such as NeurIPS, ICML, ACL, or ICLR, particularly in evaluation, benchmarking, or data-centric AI.
- Experience authoring a widely adopted public benchmark or dataset.
- Industry experience on an evaluation, benchmarking, or post-training team at a frontier lab.
- Domain expertise in finance, law, consulting, accounting, medicine, or software engineering.
- Experience designing rubrics, model-as-judge pipelines, or large-scale human annotation programs.
Compensation and Benefits
- Annual salary range: $200,000–$400,000.
- Bi-annual performance bonus structure.
- Generous equity grant vested over four years.
- Up to $15,000 relocation bonus.
- $10,000 housing bonus for employees living within 0.5 miles of the office.
- $1,500 monthly meal stipend.
- Equinox membership.
- $200 monthly laundry reimbursement.
- $200 monthly personal wellness reimbursement.
- Health, dental, and vision insurance.
- 401(k) with company match.
Skills
Llm Evaluation, Benchmarking, Natural Language Processing, Experimental Design, Statistics, Python, Evaluation Harnesses, Contamination Controls, Inter-Rater Agreement, Model-As-Judge, Human Annotation, Data-Centric Ai, Machine Learning, Research Publishing
Similar jobs
AI Research jobsResearch Scientist defining and executing research on reliable long-horizon agents in enterprise environments. The role focuses on post-training and reinforcement learning, agent memory, evaluation, verification, and structured representations, combining hands-on experimentation with product delivery and publication.
Conduct rigorous people research and applied data science to evaluate talent programs, organizational health, and employee experiences. The role requires advanced expertise in research design, experimentation, measurement, causal inference, statistical modeling, and responsible handling of sensitive employee data.
Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.
Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.
Research novel post-training methods for large language models, focusing on preference optimization, data curation, evaluation, alignment, and robustness across text and multimodal systems. Requires advanced academic training and experience with deep learning, reinforcement learning, and post-training techniques.