# Product Designer, Evals & Prompts

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** San Francisco, CA
**Role:** Product Design
**Salary:** $305k – $385k/yr
**Skills:** Python, Llm Evaluations, Prompt Engineering, Evaluation Pipelines, Graders, Regression Testing, Test Harnesses, Sandboxing, A/B Testing, Front-End Development, Human Feedback, Preference Pairs
**Posted:** 2026-09-04

> Build LLM product evaluations, prompt improvements, and designer-facing tools that measure and improve Claude’s behavior across product surfaces and model releases. The role requires production Python, evaluation-pipeline experience, internal tool development, and prompt or model-launch experience.

## Job Description

## Responsibilities
- Write, test, revise, and ship prompts for Claude’s tools, features, and product behaviors; verify that shipped prompts match the intended versions.
- Build graders, automated evaluations, comparison sets, regression suites, and the infrastructure that runs them across models.
- Develop visual, low-code evaluation tools that enable designers to assemble transcript sets, create graders from plain-English rubrics, compare prompt variants, and inspect results.
- Observe designers using evaluation tools and simplify the workflows.
- Support model releases by testing product surfaces, creating prompt fixes and migrations, and developing prompts for new features.
- Build and scale a reliable evaluation harness for tool calls, including sandboxing, consistent settings, and regression diagnosis.
- Package behaviors for training using graders, human-feedback questions, and preference pairs.

## Requirements
- Production-quality Python.
- Experience building and maintaining evaluation pipelines for LLM products, including graders, rubrics, comparison sets, regression suites, and model-run infrastructure.
- Experience building internal tools with interfaces for nontechnical users.
- Experience standing up test harnesses, sandboxing tool calls, and pinning comparable run settings.
- Experience shipping prompts or collaborating closely with prompt engineers, with an understanding of cross-model prompt behavior.
- Ability to analyze transcripts in addition to evaluation scores.
- Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.

## Nice-to-haves
- Experience working within model-launch cycles.
- A/B testing experience and ability to connect offline evaluations to online outcomes.
- Front-end development or notebook-to-application experience, with strong opinions about evaluation-result usability.
- Experience turning product rubrics into training signals, including graders, human-feedback questions, or preference pairs.
- Product-focused approach to model behavior and user experience.

## Compensation
- Annual salary: **$305,000–$385,000 USD**.

## Similar jobs

- [Product Design Leadership, Growth](https://hotfix.jobs/jobs/dd25af36-0122-43d3-83c6-865275b769af) - OpenAI - San Francisco, CA - $347k – $405k/yr
- [Product Designer, Identity](https://hotfix.jobs/jobs/a6d6dc1c-a7c4-47a9-93fc-192a623c5457) - OpenAI - San Francisco, CA - $245k – $310k/yr
- [Product Designer](https://hotfix.jobs/jobs/1a1369c8-102c-40d2-b6cb-854984e48cc1) - Baseten - San Francisco, CA - $225k – $290k/yr
- [Product Designer](https://hotfix.jobs/jobs/1b3ef200-5cb1-4c80-b05e-2c90580732b6) - F2 - San Francisco, CA - $220k – $250k/yr
- [Web Designer](https://hotfix.jobs/jobs/e1fb7519-553b-4fef-b52b-906c662800bb) - OpenAI - New York, NY - $216k – $240k/yr

**Apply:** https://hotfix.jobs/jobs/1b36c6bb-d7f6-4a2a-aa5c-0cfb1c829b02
**Canonical:** https://hotfix.jobs/jobs/1b36c6bb-d7f6-4a2a-aa5c-0cfb1c829b02