Skip to content
ZooxZooxFoster City, CA

Software Development Engineer in Test, Machine Learning

Build and productionize agentic LLM-powered triage systems and ML/DL pipelines to automate failure analysis for autonomous robots. Requires Master's/PhD in STEM + 4+ years production ML/NLP experience with PyTorch, RAG, Databricks, and AWS.

180k – 225k/yr
Hybrid4+ YOEML Engineering

About the role

Responsibilities

  • Design and improve ML/DL algorithms on large-scale data to automate test and triage workflows.
  • Build and run agentic LLM systems (e.g., Claude or Gemini) that automate triage, from prototype to production.
  • Run multiple production pipelines: monitor them, respond to issues, and fix the root causes.
  • Work with stakeholders across data science, ML, autonomous drive planning, and quality assurance.
  • Build monitoring tools and use them to keep improving your algorithms in production.

Requirements

  • PhD or Master's in a STEM field and 4+ years in Triage, Performance Analytics for Fortune 500, including infrastructure as code, ML, NLP, or deep learning.
  • Strong productionization and CI/CD experience for medium to large projects in Python, with a solid grasp of dashboards, eval, algorithms, and data manipulation.
  • Experience with RAG, PyTorch, TensorFlow, or HuggingFace.
  • Has built or managed agentic LLM systems (Claude or Gemini) in a professional setting with large datasets to satisfy multiple stakeholders.
  • Has run multiple production pipelines on platforms like Databricks and AWS, including fixing production issues.

Nice-to-Haves

  • Experience with self-correcting models, installation and maintenance of LLMs, data visualization tools like Looker/Dash/Databricks Dashboards.
  • Experience in Performance Monitoring at Scale, Support environments, autonomous driving, robotics.
  • Strong verbal and written communication skills.

Skills

PythonPyTorchTensorFlowhuggingfaceRAGLLMsAWSDatabricksCI/CDMachine LearningDeep LearningNLP

Similar roles

ML Engineering jobs
Applied Intuition

Perception Engineer

Applied IntuitionSunnyvale, CA

Perception Engineer owning outcomes for autonomous mining vehicles. Responsible for sensor selection, model adaptation to new sites/domains, diagnosing failures, data strategies, and translating customer needs into technical KPIs and solutions. Requires strong systems understanding of perception/full stack and real-world deployment experience.

180k – 255k/yr
On-site5+ YOEML Engineering
Sprinter Health

Applied Scientist, AI

Sprinter HealthSan Francisco, CA

Build, evaluate, and productionize ML/AI models (including LLMs and NLP) that solve ambiguous healthcare, product, and operational problems at Sprinter Health. Requires strong experimentation, error analysis, stakeholder collaboration with clinicians, and focus on real-world impact, bias, and evaluation.

180k – 260k/yr
HybridML Engineering
Notion

Software Engineer, AI Platform

NotionSan Francisco, CA +1

Build and scale the shared AI platform foundations at Notion, enabling fast and safe shipping of AI products. Requires experience with LLM/ML platforms, strong ownership, and comfort across backend, infrastructure, and product code.

180k – 201k/yr
Hybrid5+ YOEML Engineering
Exa

Research Engineer, Generalist

ExaSan Francisco, CA

Generalist Research Engineer working across Exa's search and retrieval stack including crawling, parsing, ML performance, and retrieval algorithms to improve search quality and performance for customers.

180k – 350k/yr
On-siteML Engineering
Baseten

Software Engineer - BIS

BasetenSan Francisco, CA

As a Software Engineer on the Inference Stack team, you will build the distributed runtime that powers large-scale LLM inference. This role involves working across the stack, from developer experience to low-level infrastructure, and owning systems in production.

180k – 360k/yr
HybridML Engineering