Skip to content
ArenaArena

Machine Learning Scientist - Open Source Lead

Leads open-source ML research by designing experiments, developing evaluation methodologies, analyzing preference data, and releasing datasets/code to advance AI model transparency. Requires PhD-level expertise in ML/LLMs and hands-on experience with RLHF/DPO fine-tuning.

About the job

Responsibilities

  • Design and conduct experiments to evaluate AI model behavior across reasoning, style, robustness, and user preference dimensions
  • Develop new metrics, methodologies, and evaluation protocols that go beyond traditional benchmarks
  • Analyze large-scale human voting and interaction data to uncover insights into model performance and user preferences
  • Communicate results with the broader research community via academic papers, educational content, conference talks
  • Collaborate with engineers to implement and scale research findings into production systems
  • Prototype and test research ideas rapidly, balancing rigor with iteration speed
  • Partner with model providers to shape evaluation questions and support responsible model testing
  • Contribute to the scientific integrity and transparency of the LMArena leaderboard and tools

Requirements

  • Hands-on experience training large-scale models, including reward models, preference models, and fine-tuning LLMs with methods like RLHF, DPO, and contrastive learning
  • Strong foundation in ML and statistics, with a track record of designing novel training objectives, evaluation schemes, or statistical frameworks to improve model reliability and alignment
  • Fluent in the full experimental stack, from dataset design and large-batch training to rigorous evaluation and ablation, with an eye for what scales to production
  • Deeply collaborative mindset, working closely with engineers to productionize research insights and iterating with product teams to align research with user needs
  • Comfortable being a visible representative of Arena Intelligence, engaging openly with the research community, and building a strong personal brand to help shape AI research culture
  • PhD or equivalent research experience in Machine Learning, Natural Language Processing, Statistics, or a related field
  • Strong understanding of LLMs and modern deep learning architectures (e.g., Transformers, diffusion models, reinforcement learning with human feedback)
  • Proficiency in Python and ML research libraries such as PyTorch, JAX, or TensorFlow
  • Demonstrated ability to design and analyze experiments with statistical rigor
  • Experience publishing research or working on open-source projects in ML, NLP, or AI evaluation
  • Comfortable working with real-world usage data and designing metrics beyond standard benchmarks
  • Ability to translate research questions into practical systems and collaborate across engineering and product teams
  • Passion for open science, reproducibility, and community-driven research

Nice-to-Haves

  • Skilled at public speaking, writing, and presenting research work to diverse audiences
  • Actively participates in conferences, panels, and online forums to foster relationships and thought leadership
  • Builds trust through transparent communication and consistent community engagement
  • Serves as a go-to contact for external researchers, journalists, and partners

Skills

PyTorch, JAX, TensorFlow, Python, LLMs, Transformers, RLHF, Dpo, Machine Learning, Statistics

Decagon

Decagon

San Francisco, CA
Senior Research Engineer, Safety
$200k+/yrOn-site4+ YOEAI Research

Research and build safety models, evaluations, and runtime safeguards for conversational AI agents, addressing prompt injection, unsafe tool use, privacy, and policy risks. Requires 4+ years in AI/ML engineering, research, or safety plus experience deploying and evaluating language models or agentic systems.

Function Health

Function Health

United States

Senior Clinical Specialist, AI Systems
No salary listedRemoteAI Research

Evaluates and improves AI-generated clinical outputs, partnering with product and engineering teams to establish safety, accuracy, and clinical-quality standards. Requires an MD, DO, or equivalent clinical doctorate, substantial patient-care experience, strong clinical judgment, and the ability to learn AI evaluation techniques.

Gusto

Gusto

Denver, CO
AI Solutions Architect
$168k+/yrHybrid8+ YOEAI Research

Owns reusable patterns, standards, and tooling for production agentic service workflows, guiding platform priorities, automation measurement, and quality governance. Requires 8+ years in operations, product, or AI, hands-on agentic workflow experience, and strong LLM, metrics, and cross-functional influence skills.

Anthropic

Anthropic

San Francisco, CA
Applied AI, Research Engineer
$300k+/yrHybrid6+ YOEAI Research

Applied AI Research Engineer who tests model capabilities, builds demos and evaluations, supports strategic customer implementations, and translates field insights into product and research direction. Requires 6+ years of technical experience, programming proficiency, LLM development experience, and strong communication skills.

Ai2

Ai2

Seattle, WA

Senior Research Scientist, Open Ecosystem
$170k+/yrOn-site5+ YOEAI Research

Conduct research and build open foundation models and training systems aimed at accelerating scientific discovery. The role requires a PhD-level background and substantial experience training foundation models, with expertise in agentic training or multimodal data preferred.