Skip to content
Distyl AIDistyl AI

AI Engineer, Evaluation

Design and implement evaluation frameworks and pipelines for AI systems using Evaluation-Driven Development. Build Python-based test suites, LLM graders, and measurement systems that guide prompt iteration and production deployment decisions.

About the job

Key Responsibilities

  • Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments
  • Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives
  • Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases
  • Develop and maintain evaluation pipelines—offline and online—that integrate directly into system iteration loops
  • Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments
  • Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts

What We Require

  • 2+ years of software engineering experience
  • Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments
  • Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration
  • Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders
  • Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment
  • AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration
  • Travel: Travel between 10-50% of the time, depending on the project

What We Offer

  • Base salary range: $150K – $250K
  • Meaningful equity
  • 100% covered medical, dental, and vision for employees and dependents
  • 401(k) with additional perks
  • Access to state-of-the-art models and modern AI tools
  • Offices in San Francisco and New York with hybrid collaboration model (3+ days per week Tuesday–Thursday in-office)

Skills

Python, Evaluation Frameworks, Experimentation, Llm-Based Graders, Prompt Engineering, Ai Systems, Test Case Development, Regression Testing, Production Pipelines, Model Evaluation

Garner Health

Garner Health

New York, NY

Applied Scientist II
$158k+/yrHybrid2+ YOEML Engineering

Build and ship production algorithmic systems that improve healthcare quality, access, and cost outcomes. The role combines machine learning, optimization, experimentation, and LLM productionization, requiring at least two years of relevant industry or advanced-degree experience.

Lyft

Lyft

San Francisco, CA
Machine Learning Engineer
$141k+/yrHybrid2+ YOEML Engineering

Design, deploy, and improve real-time machine learning systems for Lyft’s ride fulfillment and marketplace products. The role requires 2+ years of ML experience, production programming skills, and expertise with deep learning and recommendation systems.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff
$160k+/yrOn-siteML Engineering

Build, deploy, and optimize AI applications and machine learning models for customer use cases while contributing to an internal ML platform. This new graduate role requires a technical master’s degree, hands-on ML or LLM experience, and strong customer communication skills.

Ambral

Ambral

New York, NY
Member of Technical Staff
$140k+/yrOn-siteML Engineering

Build production infrastructure for replayable enterprise environments, agent evaluation, and continuous model improvement. The role combines hands-on customer deployment, research experimentation, large-scale data processing, and production software engineering.

Nuro

Nuro

Mountain View, CA

Software Engineer, ML Inference Platform
$160k+/yrOn-site1+ YOEML Engineering

Build and maintain machine learning infrastructure for autonomy teams, including model pipelines, observability, inference serving, and compiler platforms. The role requires a relevant degree, at least one year of experience, strong Python skills, and familiarity with C++.