Skip to content
OtterOtterMountain View, CA

Software Engineer, AI Agent & LLM

Build and improve AI agent experiences and conversational knowledge engine for Otter AI Chat. Focus on quality evaluation, infrastructure for orchestration/tracing, diagnosing failures across the stack, and driving measurable improvements in AI systems from traces and feedback. Requires 3+ years AI/ML engineering, strong backend/distributed systems skills, and experience shipping LLM-powered products.

155k – 185k/yr
Hybrid3+ YOEML Engineering

About the role

Responsibilities

  • Build AI-agent experiences for Otter AI Chat that help users reason over conversations, retrieve knowledge, and complete complex tasks.
  • Advance Otter’s conversational knowledge engine through better knowledge extraction, annotation, and indexing.
  • Develop robust evaluation datasets, automated graders, regression tests, and release gates for AI quality.
  • Diagnose failures across models, prompts, retrieval, tools, data pipelines, backend services, and product workflows.
  • Build shared agent infrastructure for orchestration, tracing, debugging, retries, sandboxed execution, and observability.
  • Improve task completion, correctness, groundedness, reliability, latency, and cost.
  • Turn production traces, customer feedback, and usage signals into measurable product and model improvements.
  • Collaborate across AI, product, infrastructure, data, security, and application teams to deliver end-to-end capabilities.

Requirements

  • 3+ years of AI Agent engineering, machine learning engineering, or related experience.
  • Strong backend or distributed-systems engineering skills.
  • Hands-on experience building and shipping AI-agent or LLM-powered products.
  • Strong focus on AI quality and experience evaluating nondeterministic systems.
  • Ability to use data, traces, logs, and qualitative examples to identify and resolve complex failures.
  • Works effectively across services, technical domains, and organizational boundaries.
  • Combines strong product judgment with rigorous engineering and evaluation practices.
  • High agency, strong ownership, and a bias toward action.
  • Ability to take ambiguous problems from initial exploration through production launch and continuous improvement.

Nice-to-Haves

  • Demonstrated ability to use coding agents effectively while rigorously reviewing and controlling the quality of their output.

Skills

AI AgentsLLMsMachine Learningbackend engineeringDistributed Systemsevaluation systemsPrompt EngineeringData PipelinesObservabilityPython

Similar roles

ML Engineering jobs
Ramp

Applied AI Engineer

RampNew York, NY +1

Build and ship full-stack AI projects including AI agents, RAG, structured extraction, and LLM infrastructure. Requires proficiency in full-stack development, backend systems, cloud infrastructure, and production LLM experience.

155k – 340k/yr
HybridML Engineering
Twilio

Machine Learning Engineer

TwilioUnited States

Build, deploy, and operate scalable ML systems for real-time applications like anomaly detection, recommendations, and agentic AI at Twilio. Requires 5+ years production ML experience, strong Python/Java/SQL skills, and MLOps expertise.

156k – 229k/yr
Remote5+ YOEML Engineering
Applied Intuition

Software Engineer - E2E Autonomy

Applied IntuitionSunnyvale, CA

Builds ML tools, infrastructure, and manages large datasets for end-to-end autonomy research and productionizing self-driving software. Works with AI research and engineering teams to scale GPU compute, data, and evaluation systems. Requires strong software generalist skills across ML stack.

153k – 222k/yr
On-siteML Engineering
Writer

AI engineer

WriterSan Francisco, CA +1

Build and deploy scalable AI applications and intelligent agents for enterprise customers. Requires 5+ years of experience with Python, ML frameworks (PyTorch/TensorFlow/JAX), LLMs, and cloud platforms.

152k – 316k/yr
Hybrid5+ YOEML Engineering
Brex

Software Engineer, Forward Deployed Agent Builder

BrexSeattle, WA +2

Builds and deploys AI agents to automate internal workflows at Brex by embedding with teams, integrating systems and APIs, and measuring impact. Requires 4+ years experience shipping AI/automation with LLMs, agent frameworks, and databases.

152k – 240k/yr
Hybrid4+ YOEML Engineering