Technical Lead, Evaluation Infrastructure
Lead the Evaluation Infrastructure team building metrics, evaluation pipelines, and validation platforms for autonomous vehicle safety and iteration. Requires 4+ years experience in distributed systems and ML evaluation, plus strong Python/C++ skills and AI-native engineering practices.
About the job
Responsibilities
- Build and own a unified metrics, evaluation, and validation platform — pipelines, introspection tooling, and analysis products that turn on-road and simulation logs into high-fidelity signals for autonomy iteration and driverless safety validation
- Drive the technical bar for metric quality across both heuristic and ML-based approaches
- Invest in the scale, reliability, and CI/CD of the evaluation stack to shorten time-to-signal for evaluation and time-to-confidence for validation, and to meet high SLAs for downstream stakeholders
- Mentor and grow the Evaluation Infrastructure team, and champion AI-native engineering practices that compound team velocity and code quality
- Partner with Product, Autonomy, Systems & Safety, and Simulation teams to define and execute the vision and strategy for evaluation
Requirements
- B.Sc or M.Sc. degree plus 4 years of relevant work experience
- Strong fluency in distributed systems, large-scale data and ML evaluation pipelines, metrics frameworks (heuristic and/or ML-based), and analytics platforms
- Experience setting technical vision, roadmap, and prioritization for a team operating at the intersection of autonomy, safety, and data infrastructure
- Clear, concise communicator who partners effectively with PMs, engineers, and cross-functional stakeholders
- Ability and willingness to deep-dive into implementation
- Sets the technical bar for metric quality, pipeline rigor, and safety-critical engineering practice
- Strong proficiency in Python, C++, or similar languages
- Daily user of modern AI coding assistants and agentic tools (Claude Code, Cursor, and similar), with strong intuition for where they accelerate engineering work
Nice-to-Haves
- Knowledge of data engineering tooling and best practices
- Knowledge of batch and streaming data processing, warehousing, and analytics solutions
- Experience with data workflow orchestration platforms
- Prior experience building evaluation, validation, or analytics platforms, ideally in autonomy, robotics, or safety-critical systems
Compensation & Benefits
- Base pay range: $193,930 - $291,150/year
- Annual performance bonus and equity
- Competitive benefits package
Skills
Python, C++, Distributed Systems, Ml Evaluation Pipelines, Metrics Frameworks, Data Engineering, Batch Processing, Streaming Data, Data Workflow Orchestration, CI/CD
Similar jobs
ML Engineering jobsBuild and operate large-scale infrastructure for autonomous-driving model training, including distributed GPU systems, data pipelines, ML workflows, and reliability tooling. The role requires 3+ years of experience, strong Python and systems-language skills, Kubernetes expertise, and distributed-systems fundamentals.
Leads the design, deployment, and optimization of agentic and generative AI systems that automate risk and compliance investigations at scale. Requires 8+ years of machine learning modeling experience, production ML expertise, and advanced technical education.
Design and deploy tactical autonomy algorithms and high-performance software for unmanned systems operating in complex, contested environments. The role requires 5+ years of related experience, strong C++ and Python skills, robotics expertise, and the ability to obtain a SECRET clearance.
Senior machine learning engineer who will build and operate large-scale AI systems for Airbnb’s payments ecosystem, including LLM agents, fraud defenses, and personalization. The role requires 5+ years of applied AI/ML experience, strong Python or Java skills, and production MLOps expertise.
Build and operate large-scale machine learning infrastructure and models for Reddit’s recommendation and personalization systems. The role requires 5+ years of ML engineering experience, expertise in deep learning and distributed systems, and proficiency with Python and modern ML frameworks.