Machine Learning Scientist - Open Source Lead
Leads open-source ML research by designing experiments, developing evaluation methodologies, analyzing preference data, and releasing datasets/code to advance AI model transparency. Requires PhD-level expertise in ML/LLMs and hands-on experience with RLHF/DPO fine-tuning.
About the job
Responsibilities
- Design and conduct experiments to evaluate AI model behavior across reasoning, style, robustness, and user preference dimensions
- Develop new metrics, methodologies, and evaluation protocols that go beyond traditional benchmarks
- Analyze large-scale human voting and interaction data to uncover insights into model performance and user preferences
- Communicate results with the broader research community via academic papers, educational content, conference talks
- Collaborate with engineers to implement and scale research findings into production systems
- Prototype and test research ideas rapidly, balancing rigor with iteration speed
- Partner with model providers to shape evaluation questions and support responsible model testing
- Contribute to the scientific integrity and transparency of the LMArena leaderboard and tools
Requirements
- Hands-on experience training large-scale models, including reward models, preference models, and fine-tuning LLMs with methods like RLHF, DPO, and contrastive learning
- Strong foundation in ML and statistics, with a track record of designing novel training objectives, evaluation schemes, or statistical frameworks to improve model reliability and alignment
- Fluent in the full experimental stack, from dataset design and large-batch training to rigorous evaluation and ablation, with an eye for what scales to production
- Deeply collaborative mindset, working closely with engineers to productionize research insights and iterating with product teams to align research with user needs
- Comfortable being a visible representative of Arena Intelligence, engaging openly with the research community, and building a strong personal brand to help shape AI research culture
- PhD or equivalent research experience in Machine Learning, Natural Language Processing, Statistics, or a related field
- Strong understanding of LLMs and modern deep learning architectures (e.g., Transformers, diffusion models, reinforcement learning with human feedback)
- Proficiency in Python and ML research libraries such as PyTorch, JAX, or TensorFlow
- Demonstrated ability to design and analyze experiments with statistical rigor
- Experience publishing research or working on open-source projects in ML, NLP, or AI evaluation
- Comfortable working with real-world usage data and designing metrics beyond standard benchmarks
- Ability to translate research questions into practical systems and collaborate across engineering and product teams
- Passion for open science, reproducibility, and community-driven research
Nice-to-Haves
- Skilled at public speaking, writing, and presenting research work to diverse audiences
- Actively participates in conferences, panels, and online forums to foster relationships and thought leadership
- Builds trust through transparent communication and consistent community engagement
- Serves as a go-to contact for external researchers, journalists, and partners
Skills
PyTorch, JAX, TensorFlow, Python, LLMs, Transformers, RLHF, Dpo, Machine Learning, Statistics
Similar jobs
AI Research jobsResearch and build safety models, evaluations, and runtime safeguards for conversational AI agents, addressing prompt injection, unsafe tool use, privacy, and policy risks. Requires 4+ years in AI/ML engineering, research, or safety plus experience deploying and evaluating language models or agentic systems.
Evaluates and improves AI-generated clinical outputs, partnering with product and engineering teams to establish safety, accuracy, and clinical-quality standards. Requires an MD, DO, or equivalent clinical doctorate, substantial patient-care experience, strong clinical judgment, and the ability to learn AI evaluation techniques.
Owns reusable patterns, standards, and tooling for production agentic service workflows, guiding platform priorities, automation measurement, and quality governance. Requires 8+ years in operations, product, or AI, hands-on agentic workflow experience, and strong LLM, metrics, and cross-functional influence skills.
Applied AI Research Engineer who tests model capabilities, builds demos and evaluations, supports strategic customer implementations, and translates field insights into product and research direction. Requires 6+ years of technical experience, programming proficiency, LLM development experience, and strong communication skills.
Conduct research and build open foundation models and training systems aimed at accelerating scientific discovery. The role requires a PhD-level background and substantial experience training foundation models, with expertise in agentic training or multimodal data preferred.