Skip to content

Software Engineer

Builds and debugs AI agent infrastructure for healthcare automation, including prompt engineering, runtime issue tracing, evaluation datasets, simulation tooling, data pipelines, and observability dashboards. Requires 2-7 years experience with production LLMs/AI agents and TypeScript proficiency.

About the job

What You’ll Do

  • Work across our AI agent platform — writing prompts, debugging runtime issues, building agent simulation tooling, creating evals, interfacing with client data, and helping monitor system behavior at scale.
  • Trace and fix runtime bugs, then write regression tests.
  • Design evaluation datasets to simulate realistic workflows or red-team our system.
  • Build internal tooling for QA and agent simulation.
  • Normalize and transform messy client data for system integration.
  • Set up automatic testing and latency tracking infrastructure.
  • Create dashboards and observability tooling for agentic system behavior.
  • Expand on our existing eval & testing framework and agent simulation infrastructure.

Skills Required

Technical Skills

  • Proficiency in TypeScript
  • Strong generalist software engineering
  • Strong debugging skills (trace runtime failures, dig through logs, pinpoint issues in async or multi-step agent systems)
  • Data transformation and ingestion (build pipelines to normalize and convert unstructured data for AI systems)
  • Strong understanding of system design, including distributed systems and reliability/performance tradeoffs
  • Experience using modern AI coding tools (e.g. Cursor, GitHub Copilot, Claude)
  • Excellent documentation and testing discipline
  • Proficiency with Git

Soft Skills

  • Care about improving agent behavior
  • High agency; thrive with minimal structure
  • Comfortable getting in the weeds with details, edge cases, editing prompts, writing evals
  • Comfortable with ambiguity; work well with loose specs spanning prompts, code, RLHF
  • Learn fast and move fast; pattern-match from past systems work to LLM edge cases

Experience & Who Should Apply

  • 2-7 years of experience working closely with LLMs or AI agents in production systems
  • Created internal tools or frameworks for QA, evals, or agent simulation
  • Contributed to fast-paced product cycles involving AI behavior, latency, user experience

Nice to Have

  • Experience with multi-agent systems, TTS/NLP pipelines, or structured output validation
  • Familiarity with testing frameworks, LangChain-style agent orchestration, or in-house eval harnesses
  • Experience with prompt engineering, LLM evals, and agent orchestration

Skills

TypeScript, LLMs, AI Agents, Git, Cursor, Github Copilot, Claude, LangChain, Prompt Engineering, Llm Evals, Agent Orchestration, Distributed Systems

Ambral

Ambral

New York, NY
Member of Technical Staff
$140k+/yrOn-siteML Engineering

Build production infrastructure for replayable enterprise environments, agent evaluation, and continuous model improvement. The role combines hands-on customer deployment, research experimentation, large-scale data processing, and production software engineering.

Lyft

Lyft

San Francisco, CA
Machine Learning Engineer
$141k+/yrHybrid2+ YOEML Engineering

Design, deploy, and improve real-time machine learning systems for Lyft’s ride fulfillment and marketplace products. The role requires 2+ years of ML experience, production programming skills, and expertise with deep learning and recommendation systems.

Pinterest

Pinterest

San Francisco, CA

Machine Learning Engineer II, Responsible AI
$139k+/yrHybrid2+ YOEML Engineering

Develop and deploy responsible AI and machine learning fairness solutions across Pinterest’s large-scale, user-facing products, including generative AI, search, and recommendations. The role requires production ML experience, expertise in fairness interventions and modern architectures, and a master’s or PhD in computer science or a related field.

Benchling

Benchling

San Francisco, CA

Software Engineer, Model Evaluation and Improvement
$136k+/yrOn-site2+ YOEML Engineering

Build datasets, evaluations, and scalable data systems that improve frontier AI models on challenging biological and scientific tasks. The role partners with scientists and AI labs and requires at least two years of experience applying biology and AI, plus hands-on LLM experience.

DataVisor

DataVisor

Mountain View, CA

Software Engineer, Artificial Intelligence
$130k+/yrOn-site2+ YOEML Engineering

Builds high-scale data pipelines, distributed systems, and AI agent workflows using LLMs for fraud intelligence platform. Requires 2+ years software engineering, Python proficiency, big data tools, AWS/K8s, and ML foundations.