Senior Research Scientist, Model Evaluation
Conducts research and builds benchmarks, methods, and infrastructure to evaluate the capabilities and progress of large language models. The role requires strong software engineering, prototyping, data-quality review, and rigorous measurement skills.
About the job
Responsibilities
- Create ambitious new evaluation benchmarks that push the limits of what models can accomplish.
- Work on highly cross-functional teams to translate model feedback into trustworthy, repeatable evaluations.
- Conduct research to advance the state of the art in LLM evaluation methods, including:
- Training LLM judges.
- Refining LLM-based data synthesis pipelines.
- Improving evaluation efficiency.
- Build scalable and reusable tools for analyzing model performance.
Requirements and Qualifications
- Experience rapidly building prototypes that demonstrate the boundaries of LLM capabilities.
- Experience developing resources to measure model capabilities.
- Extensive experience reviewing complex data and LLM outputs to ensure high data quality.
- Strong focus on rigorously measuring AI capabilities and aligning measurements with the capabilities being evaluated.
- Strong software engineering skills.
Compensation and Benefits
- Weekly lunch stipend of $75, £75, or equivalent in local currency.
- Full health and dental benefits, including a separate mental health budget.
- RRSP matching, 401(k), or pension scheme.
- 100% parental leave top-up for up to six months for either parent.
- Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation.
- Travel budget for remote employees visiting other offices and an annual company offsite.
- Coworking benefit for employees who are not near an office.
- $500 home office stipend.
Skills
LLMs, Llm Evaluation, Evaluation Benchmarks, Llm Judges, Data Synthesis, Software Engineering, Prototyping, Data Quality, Ai Capabilities, Model Performance
Similar jobs
AI Research jobsResearch and build safety models, evaluations, and runtime safeguards for conversational AI agents, addressing prompt injection, unsafe tool use, privacy, and policy risks. Requires 4+ years in AI/ML engineering, research, or safety plus experience deploying and evaluating language models or agentic systems.
Evaluates and improves AI-generated clinical outputs, partnering with product and engineering teams to establish safety, accuracy, and clinical-quality standards. Requires an MD, DO, or equivalent clinical doctorate, substantial patient-care experience, strong clinical judgment, and the ability to learn AI evaluation techniques.
Owns reusable patterns, standards, and tooling for production agentic service workflows, guiding platform priorities, automation measurement, and quality governance. Requires 8+ years in operations, product, or AI, hands-on agentic workflow experience, and strong LLM, metrics, and cross-functional influence skills.
Applied AI Research Engineer who tests model capabilities, builds demos and evaluations, supports strategic customer implementations, and translates field insights into product and research direction. Requires 6+ years of technical experience, programming proficiency, LLM development experience, and strong communication skills.
Conduct research and build open foundation models and training systems aimed at accelerating scientific discovery. The role requires a PhD-level background and substantial experience training foundation models, with expertise in agentic training or multimodal data preferred.