Research Scientist
Conduct original research on LLM evaluation, routing optimization, and model behavior using billions of real-world generations. Design novel benchmarks, run large-scale empirical studies, and develop statistical foundations for intelligent routing. Requires MS/PhD, publication track record, deep stats/ML expertise, and Python/SQL skills.
About the job
What You'll Do
- Own and pursue a research agenda focused on LLM evaluation, model quality, routing optimization, and AI usage patterns, contributing original insights that advance the field.
- Design novel evaluation frameworks and benchmarks that go beyond standard leaderboards, using real-world generation data to capture how models actually perform across tasks and contexts.
- Conduct large-scale empirical studies on LLM behavior: how models compare across providers, how performance changes over time, and how usage patterns reveal strengths and weaknesses.
- Develop the statistical and mathematical foundations behind our routing systems, building the models and heuristics that power intelligent provider and model selection.
- Identify opportunities to apply research findings to feed back into OpenRouter's product and platform.
- Collaborate with external researchers, model providers, and the open-source community to advance shared understanding of LLM capabilities and limitations.
- Work with product and engineering teams to translate research findings into improvements to OpenRouter's platform.
What You Bring
Experience & Technical Skills
- MS or PhD in a quantitative field (machine learning, statistics, computer science, mathematics, computational linguistics, or similar).
- Track record of original research, demonstrated by first-author publications, significant open-source contributions, or equivalent impact in industry research.
- Deep expertise in statistics, experimental design, and causal inference.
- Strong programming skills in Python for building data pipelines, running large-scale experiments, and prototyping models.
- Proficiency in SQL for working with large-scale analytical databases (ClickHouse, BigQuery, or similar).
- Hands-on experience with modern ML/NLP techniques such as LLM evaluation, fine-tuning, embeddings, classification, or reinforcement learning from human feedback.
- Familiarity with the current LLM landscape: model architectures, provider ecosystems, benchmark suites, and the strengths and limitations of leading models.
Mindset & Approach
- Deeply curious and self-directed.
- Rigorous but pragmatic, operating at startup speed while maintaining high scientific standards.
- AI-first in your own workflow, using LLMs, coding agents, and modern AI tools heavily.
- Strong communicator, able to explain complex findings clearly.
- Collaborative with product and engineering teams.
Skills
Python, SQL, Machine Learning, Statistics, Causal Inference, Llm Evaluation, Fine-Tuning, Embeddings, ClickHouse, BigQuery
Similar jobs
AI Research jobsConduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.
Research Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.
Conduct applied research on foundation models for fraud detection using large-scale behavioral and financial-risk data. The role spans experimentation, evaluation, production deployment, and cross-functional work on model governance, requiring 4+ years of applied ML experience and strong Python and SQL skills.
Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.
Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.