AI Engineer, Evaluation
Design and implement evaluation frameworks and pipelines for AI systems using Evaluation-Driven Development. Build Python-based test suites, LLM graders, and measurement systems that guide prompt iteration and production deployment decisions.
About the job
Key Responsibilities
- Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments
- Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives
- Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases
- Develop and maintain evaluation pipelines—offline and online—that integrate directly into system iteration loops
- Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments
- Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts
What We Require
- 2+ years of software engineering experience
- Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments
- Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration
- Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders
- Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment
- AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration
- Travel: Travel between 10-50% of the time, depending on the project
What We Offer
- Base salary range: $150K – $250K
- Meaningful equity
- 100% covered medical, dental, and vision for employees and dependents
- 401(k) with additional perks
- Access to state-of-the-art models and modern AI tools
- Offices in San Francisco and New York with hybrid collaboration model (3+ days per week Tuesday–Thursday in-office)
Skills
Python, Evaluation Frameworks, Experimentation, Llm-Based Graders, Prompt Engineering, Ai Systems, Test Case Development, Regression Testing, Production Pipelines, Model Evaluation
Similar jobs
ML Engineering jobsBuild and ship production algorithmic systems that improve healthcare quality, access, and cost outcomes. The role combines machine learning, optimization, experimentation, and LLM productionization, requiring at least two years of relevant industry or advanced-degree experience.
Design, deploy, and improve real-time machine learning systems for Lyft’s ride fulfillment and marketplace products. The role requires 2+ years of ML experience, production programming skills, and expertise with deep learning and recommendation systems.
Build, deploy, and optimize AI applications and machine learning models for customer use cases while contributing to an internal ML platform. This new graduate role requires a technical master’s degree, hands-on ML or LLM experience, and strong customer communication skills.
Build production infrastructure for replayable enterprise environments, agent evaluation, and continuous model improvement. The role combines hands-on customer deployment, research experimentation, large-scale data processing, and production software engineering.
Build and maintain machine learning infrastructure for autonomy teams, including model pipelines, observability, inference serving, and compiler platforms. The role requires a relevant degree, at least one year of experience, strong Python skills, and familiarity with C++.