What you’ll work on
- Advance the agent harness by bringing the latest research and open-source developments into production and experimenting with new approaches to multi-agent coordination, model routing, memory, planning, and tool use.
- Build rigorous evaluations that measure agent quality, reliability, latency, and cost on representative scientific tasks.
- Monitor agent quality in production, troubleshoot failures quickly in high-ambiguity situations, and translate findings into engineering improvements.
- Collaborate on reliable infrastructure for long-running sessions, safe sandboxed execution, background work, and subagents.
- Partner with scientists and engineers to evaluate and productionize new agent capabilities.
Requirements
- Strong software engineering experience building production backend systems, infrastructure, or distributed systems.
- Research or production experience in machine learning, or another quantitative discipline.
- Familiarity with LLM APIs, tool calling, agent runtimes, or workflow orchestration.
- Strong quantitative judgment and the ability to determine whether an apparent improvement is real, reproducible, and meaningful.
- Ability to move between research questions, data analysis, system design, and production implementation.
- Experience or strong interest in AI for science, scientific agents, computational research, or automated scientific discovery.
- Experience working in AI-native teams that use coding agents or automation extensively.
- High ownership, clear communication, and strong engineering judgment.
We care more about demonstrated ability than a particular credential. Relevant backgrounds may include ML or LLM engineering, academic research paired with substantial software development, scientific computing, or production systems engineering with demonstrated quantitative experience.
Nice to have
- Experience with LLM evaluations, human evaluation, model judges, replay testing, benchmark design, or experiment tracking.
- Experience applying classical machine learning methods alongside LLMs in production systems.
- Experience with task queues, event streams, Kubernetes, code sandboxes, or durable workflow systems.
- An advanced degree or equivalent experience in machine learning, computational biology, physics, applied mathematics, statistics, or a related field.