Software Engineer, AI Agent & LLM
Build and improve AI agent experiences and conversational knowledge engine for Otter AI Chat. Focus on quality evaluation, infrastructure for orchestration/tracing, diagnosing failures across the stack, and driving measurable improvements in AI systems from traces and feedback. Requires 3+ years AI/ML engineering, strong backend/distributed systems skills, and experience shipping LLM-powered products.
About the job
Responsibilities
- Build AI-agent experiences for Otter AI Chat that help users reason over conversations, retrieve knowledge, and complete complex tasks.
- Advance Otter’s conversational knowledge engine through better knowledge extraction, annotation, and indexing.
- Develop robust evaluation datasets, automated graders, regression tests, and release gates for AI quality.
- Diagnose failures across models, prompts, retrieval, tools, data pipelines, backend services, and product workflows.
- Build shared agent infrastructure for orchestration, tracing, debugging, retries, sandboxed execution, and observability.
- Improve task completion, correctness, groundedness, reliability, latency, and cost.
- Turn production traces, customer feedback, and usage signals into measurable product and model improvements.
- Collaborate across AI, product, infrastructure, data, security, and application teams to deliver end-to-end capabilities.
Requirements
- 3+ years of AI Agent engineering, machine learning engineering, or related experience.
- Strong backend or distributed-systems engineering skills.
- Hands-on experience building and shipping AI-agent or LLM-powered products.
- Strong focus on AI quality and experience evaluating nondeterministic systems.
- Ability to use data, traces, logs, and qualitative examples to identify and resolve complex failures.
- Works effectively across services, technical domains, and organizational boundaries.
- Combines strong product judgment with rigorous engineering and evaluation practices.
- High agency, strong ownership, and a bias toward action.
- Ability to take ambiguous problems from initial exploration through production launch and continuous improvement.
Nice-to-Haves
- Demonstrated ability to use coding agents effectively while rigorously reviewing and controlling the quality of their output.
Skills
AI Agents, LLMs, Machine Learning, Backend Engineering, Distributed Systems, Evaluation Systems, Prompt Engineering, Data Pipelines, Observability, Python
Similar jobs
ML Engineering jobsDevelop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.
Build and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.
Research Scientist II building and improving fraud risk models and scam detection systems using audio, behavioral, and metadata signals. Requires an advanced degree and 3+ years of applied ML experience with Python and modern ML frameworks.
Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.
Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.