Software Development Engineer in Test, Machine Learning
Build and productionize agentic LLM-powered triage systems and ML/DL pipelines to automate failure analysis for autonomous robots. Requires Master's/PhD in STEM + 4+ years production ML/NLP experience with PyTorch, RAG, Databricks, and AWS.
About the job
Responsibilities
- Design and improve ML/DL algorithms on large-scale data to automate test and triage workflows.
- Build and run agentic LLM systems (e.g., Claude or Gemini) that automate triage, from prototype to production.
- Run multiple production pipelines: monitor them, respond to issues, and fix the root causes.
- Work with stakeholders across data science, ML, autonomous drive planning, and quality assurance.
- Build monitoring tools and use them to keep improving your algorithms in production.
Requirements
- PhD or Master's in a STEM field and 4+ years in Triage, Performance Analytics for Fortune 500, including infrastructure as code, ML, NLP, or deep learning.
- Strong productionization and CI/CD experience for medium to large projects in Python, with a solid grasp of dashboards, eval, algorithms, and data manipulation.
- Experience with RAG, PyTorch, TensorFlow, or HuggingFace.
- Has built or managed agentic LLM systems (Claude or Gemini) in a professional setting with large datasets to satisfy multiple stakeholders.
- Has run multiple production pipelines on platforms like Databricks and AWS, including fixing production issues.
Nice-to-Haves
- Experience with self-correcting models, installation and maintenance of LLMs, data visualization tools like Looker/Dash/Databricks Dashboards.
- Experience in Performance Monitoring at Scale, Support environments, autonomous driving, robotics.
- Strong verbal and written communication skills.
Skills
Python, PyTorch, TensorFlow, Huggingface, RAG, LLMs, AWS, Databricks, CI/CD, Machine Learning, Deep Learning, NLP
Similar jobs
ML Engineering jobsBuild and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.
Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.
Build evaluation methods, RL environments, agent tooling, and scalable infrastructure that make subjective qualities such as design and taste measurable for frontier AI models. The role requires experience with evaluations, RL environments, ML or post-training, plus strong backend engineering skills.
Build and scale generative video and multimodal models, optimizing training and inference for efficiency, throughput, and ultra-low latency. The role requires deep learning systems expertise, strong PyTorch/CUDA experience, and the ability to move research models into production.
Build production-grade AI agents, evaluation infrastructure, and developer tooling that make AI-assisted engineering faster, safer, and reusable across teams. The role requires software engineering experience, platform or internal developer-product experience, and hands-on expertise with LLM integration and orchestration.