Software Engineer, Model Evaluation and Improvement
Build datasets, evaluations, and scalable data systems that improve frontier AI models on challenging biological and scientific tasks. The role partners with scientists and AI labs and requires at least two years of experience applying biology and AI, plus hands-on LLM experience.
About the job
Responsibilities
- Build datasets for evaluating and improving frontier models, turning complex scientific data into high-quality tasks and environments for large language models.
- Analyze model failure modes by running experiments across frontier models to identify improvement opportunities.
- Build scalable data infrastructure and pipelines to curate, transform, and validate scientific data for model evaluation and improvement.
- Collaborate with frontier AI labs on approaches for improving models on challenging scientific tasks.
- Work with scientists to translate expert judgment into problems and evaluation criteria that distinguish strong model behavior.
Requirements
- 2+ years of experience at the intersection of biology and AI, including evaluating or improving scientific models or LLMs for biological applications.
- Experience building with LLMs and understanding their capabilities and limitations.
- Curiosity about frontier AI and rapidly evolving model capabilities.
- Ability to work on ambiguous problems and adapt technical approaches as the field develops.
- Collaborative approach when working with engineers, scientists, and external research partners.
- Comfort in a fast-paced environment with shifting priorities and rapid experimentation.
Work Arrangement
- In-person role in San Francisco, Monday through Friday.
Skills
Artificial Intelligence, LLMs, Biology, Scientific Models, Model Evaluation, Model Improvement, Dataset Building, Data Infrastructure, Data Pipelines, Frontier Ai
Similar jobs
ML Engineering jobsDevelop and deploy responsible AI and machine learning fairness solutions across Pinterest’s large-scale, user-facing products, including generative AI, search, and recommendations. The role requires production ML experience, expertise in fairness interventions and modern architectures, and a master’s or PhD in computer science or a related field.
Build production infrastructure for replayable enterprise environments, agent evaluation, and continuous model improvement. The role combines hands-on customer deployment, research experimentation, large-scale data processing, and production software engineering.
Design, deploy, and improve real-time machine learning systems for Lyft’s ride fulfillment and marketplace products. The role requires 2+ years of ML experience, production programming skills, and expertise with deep learning and recommendation systems.
Builds high-scale data pipelines, distributed systems, and AI agent workflows using LLMs for fraud intelligence platform. Requires 2+ years software engineering, Python proficiency, big data tools, AWS/K8s, and ML foundations.
Build and deploy production voice AI agents for customers, creating demos, debugging edge cases, improving performance, and translating feedback into product improvements. The role combines hands-on engineering, customer engagement, and pre- and post-sales delivery.