Skip to content
PhyloPhylo

Member of Technical Staff

Build and evaluate production AI agent harness systems for biomedical discovery. Own multi-agent coordination, planning, tool use, memory, rigorous evaluations, and infrastructure for reliable scientific agents. Requires strong backend/ML engineering and quantitative judgment.

About the job

What you’ll work on

  • Advance the agent harness by bringing the latest research and open-source developments into production and experimenting with new approaches to multi-agent coordination, model routing, memory, planning, and tool use.
  • Build rigorous evaluations that measure agent quality, reliability, latency, and cost on representative scientific tasks.
  • Monitor agent quality in production, troubleshoot failures quickly in high-ambiguity situations, and translate findings into engineering improvements.
  • Collaborate on reliable infrastructure for long-running sessions, safe sandboxed execution, background work, and subagents.
  • Partner with scientists and engineers to evaluate and productionize new agent capabilities.

Requirements

  • Strong software engineering experience building production backend systems, infrastructure, or distributed systems.
  • Research or production experience in machine learning, or another quantitative discipline.
  • Familiarity with LLM APIs, tool calling, agent runtimes, or workflow orchestration.
  • Strong quantitative judgment and the ability to determine whether an apparent improvement is real, reproducible, and meaningful.
  • Ability to move between research questions, data analysis, system design, and production implementation.
  • Experience or strong interest in AI for science, scientific agents, computational research, or automated scientific discovery.
  • Experience working in AI-native teams that use coding agents or automation extensively.
  • High ownership, clear communication, and strong engineering judgment.

We care more about demonstrated ability than a particular credential. Relevant backgrounds may include ML or LLM engineering, academic research paired with substantial software development, scientific computing, or production systems engineering with demonstrated quantitative experience.

Nice to have

  • Experience with LLM evaluations, human evaluation, model judges, replay testing, benchmark design, or experiment tracking.
  • Experience applying classical machine learning methods alongside LLMs in production systems.
  • Experience with task queues, event streams, Kubernetes, code sandboxes, or durable workflow systems.
  • An advanced degree or equivalent experience in machine learning, computational biology, physics, applied mathematics, statistics, or a related field.

Skills

Python, LLM APIs, Tool Calling, Agent Runtimes, Workflow Orchestration, Kubernetes, Machine Learning, Distributed Systems, Evaluations, Benchmark Design

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

Atomicmachines

Atomicmachines

Emeryville, CA

MLOps Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.

Tessera Labs

Tessera Labs

San Jose, CA
Research Engineer
$200k+/yrOn-siteML Engineering

Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.

Tessera Labs

Tessera Labs

San Jose, CA
AI Engineer
$200k+/yrHybrid3+ YOEML Engineering

Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.

Confido

Confido

New York, NY

Applied AI/ML Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.