Agent Systems Engineer
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.
About the job
Responsibilities
- Design agent architectures for planning, reasoning, tool use, memory, and integration with external systems and data.
- Improve reliability for long-running, multi-step tasks, including failure recovery.
- Build feedback loops that enable agents to improve through real-world use.
- Develop rigorous evaluations that measure agent performance and resist gaming.
- Make practical tradeoffs among quality, latency, cost, and complexity.
Requirements
- 5+ years building production ML or backend systems.
- Experience taking LLM or agent applications from prototype to production.
- Strong understanding of agent design, including planning, reasoning, tool use, orchestration, and memory.
- Experience building evaluation systems, execution tracing, and observability for agents, with a focus on reproducibility.
- Familiarity with the OpenAI Responses API, MCP, and server-side versus client-side execution.
- Adaptability, collaboration, and willingness to develop bold ideas.
Benefits
- Flexible work with in-person collaboration in the Bay Area, a distributed global-first team, and team offsites.
- Annual travel stipend to explore a country not previously visited.
- Weekly meal allowance for takeout or grocery delivery.
- Comprehensive medical benefits and generous paid time off.
Skills
Machine Learning, Backend Systems, LLMs, Agent Architecture, Planning, Reasoning, Tool Use, Orchestration, Memory Systems, Evaluation Systems, Observability, Openai Responses Api, Mcp, Execution Tracing
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.