Skip to content
PhyloPhyloSouth San Francisco, CA

Member of Technical Staff

Build and evaluate production AI agent harness systems for biomedical discovery. Own multi-agent coordination, planning, tool use, memory, rigorous evaluations, and infrastructure for reliable scientific agents. Requires strong backend/ML engineering and quantitative judgment.

200k – 300k/yr
On-site5+ YOEML Engineering

About the role

What you’ll work on

  • Advance the agent harness by bringing the latest research and open-source developments into production and experimenting with new approaches to multi-agent coordination, model routing, memory, planning, and tool use.
  • Build rigorous evaluations that measure agent quality, reliability, latency, and cost on representative scientific tasks.
  • Monitor agent quality in production, troubleshoot failures quickly in high-ambiguity situations, and translate findings into engineering improvements.
  • Collaborate on reliable infrastructure for long-running sessions, safe sandboxed execution, background work, and subagents.
  • Partner with scientists and engineers to evaluate and productionize new agent capabilities.

Requirements

  • Strong software engineering experience building production backend systems, infrastructure, or distributed systems.
  • Research or production experience in machine learning, or another quantitative discipline.
  • Familiarity with LLM APIs, tool calling, agent runtimes, or workflow orchestration.
  • Strong quantitative judgment and the ability to determine whether an apparent improvement is real, reproducible, and meaningful.
  • Ability to move between research questions, data analysis, system design, and production implementation.
  • Experience or strong interest in AI for science, scientific agents, computational research, or automated scientific discovery.
  • Experience working in AI-native teams that use coding agents or automation extensively.
  • High ownership, clear communication, and strong engineering judgment.

We care more about demonstrated ability than a particular credential. Relevant backgrounds may include ML or LLM engineering, academic research paired with substantial software development, scientific computing, or production systems engineering with demonstrated quantitative experience.

Nice to have

  • Experience with LLM evaluations, human evaluation, model judges, replay testing, benchmark design, or experiment tracking.
  • Experience applying classical machine learning methods alongside LLMs in production systems.
  • Experience with task queues, event streams, Kubernetes, code sandboxes, or durable workflow systems.
  • An advanced degree or equivalent experience in machine learning, computational biology, physics, applied mathematics, statistics, or a related field.

Skills

PythonLLM APIstool callingagent runtimesWorkflow OrchestrationKubernetesMachine LearningDistributed Systemsevaluationsbenchmark design

Similar roles

ML Engineering jobs
Clear Street

Senior / Staff AI Platform Engineer

Clear StreetUnited States

Build and own the core AI platform powering an AI-native trading copilot. Develop high-performance Rust backend for streaming, tool execution, and safe trading actions; design robust APIs with observability and security. Requires 8+ years systems programming experience.

200k – 350k/yr
Remote8+ YOEML Engineering
Clear Street

Senior / Staff AI Model Engineer

Clear StreetUnited States

Own reliability and quality for an AI copilot in a trading platform. Design evaluation systems, benchmarks, quality gates, model improvement loops, and AI monitoring for correctness, safety, and performance in market analysis and trading workflows. Requires 8+ years production software experience and strong ML eval expertise.

200k – 350k/yr
Remote8+ YOEML Engineering
Decagon

Staff Software Engineer, Agents

DecagonSan Francisco, CA

Build and own end-to-end AI agents for enterprise customers, integrating latest text/voice models and iterating based on real-world usage. Requires 8+ years of software engineering experience with Python and TypeScript.

200k – 400k/yr
On-site8+ YOEML Engineering
Nuance Labs

Member of Technical Staff - Research Fellow

Nuance LabsSeattle, WA

3-month research fellowship for early-career researchers working on frontier Multimodal LLMs, generative modeling, and real-time audiovisual AI. Own a research problem in pretraining, post-training, RL, evaluation, or multimodal modeling. Strong PyTorch and first-author tier-1 paper required.

200k – 250k/yr
On-siteML Engineering
Nuance Labs

Member of Technical Staff — Model Optimization and Inference

Nuance LabsSeattle, WA

Early-career engineer optimizing inference for real-time multimodal AI avatars. Focus on KV cache strategies, serving frameworks, quantization, and latency reduction for LLMs and diffusion models.

200k – 300k/yr
On-siteEntry levelML Engineering