# Evals Infrastructure Tech Lead / Manager

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** San Francisco, CA
**Role:** Engineering Management
**Salary:** $500k – $850k/yr
**Experience:** 7+ years
**Skills:** Python, Rust, Distributed Systems, llm inference, llm training, Observability, sandboxing
**Posted:** 2026-07-22

> Lead the Evals Infrastructure team at Anthropic building large-scale distributed systems for model evaluation, orchestration, and trustworthy metrics that inform launch decisions for frontier AI models. Requires strong distributed systems experience, Python/Rust proficiency, and people management skills with a focus on measurement quality and AI safety.

## Job Description

## Responsibilities
- Lead the team building the distributed systems that schedule, orchestrate, and execute evals for our frontier model training
- Own eval throughput and cost: compute allocation across suites, queueing against constrained accelerator pools, caching and reuse of eval work
- Build and scale the harnesses researchers use to define, run, and iterate on evals
- Make eval results trustworthy — determinism, reproducibility, and honest uncertainty quantification on reported metrics
- Ensure eval signal reaches the dashboards and reviews where launch decisions actually get made
- Contribute directly as an engineer while managing and growing the team, prioritizing its work, and coaching your reports

## Requirements
- Led technical projects end-to-end on large-scale distributed systems, and have 1+ years managing engineers (or tech-lead-with-reports experience)
- Strong in Python and Rust
- Built high-throughput, fault-tolerant systems on cloud or on-prem accelerator fleets
- Care about measurement quality, not just pipeline uptime — you'd notice if a metric moved for the wrong reason
- Communicate well with researchers and can translate research needs into infrastructure
- Deeply interested in the transformative effects of advanced AI and committed to safe development
- Bachelor’s degree or an equivalent combination of education, training, and/or experience in a field relevant to the role

## Nice-to-Haves
- Worked on LLM inference or training infrastructure
- Experience with eval or benchmarking systems, especially agentic evals requiring sandboxed execution
- Working statistical literacy — variance, confidence intervals, sample-size sufficiency for noisy metrics
- Experience with observability and regression detection over time-series metrics

## Similar roles

- [Engineering Manager, Core Experimentation](https://hotfix.jobs/jobs/bd44b1a7-9143-47d1-b3e2-4b70037782b8) - OpenAI - Seattle, WA - $441k – $490k/yr
- [Engineering Manager, Research Data Platform](https://hotfix.jobs/jobs/9a3739e2-fd47-4263-93ac-a0bf6433f096) - Anthropic - San Francisco, CA - $405k – $850k/yr
- [Engineering Manager, Cybersecurity Products](https://hotfix.jobs/jobs/ef47f12d-53d6-4ae2-bae0-a13193257839) - Anthropic - San Francisco, CA - $405k – $485k/yr
- [Engineering Manager](https://hotfix.jobs/jobs/4a5d7681-3a8a-4c3c-a9cc-3851cddbab0e) - Thinking Machines Lab - San Francisco, CA - $400k – $500k/yr
- [Senior Enterprise Sales Manager | Housing](https://hotfix.jobs/jobs/7fb61e35-ae84-4995-a680-4e95b6cab424) - EliseAI - New York, NY - $350k – $400k/yr

**Apply:** https://hotfix.jobs/jobs/f2562691-d834-4874-af17-fa79741d5a86
**Canonical:** https://hotfix.jobs/jobs/f2562691-d834-4874-af17-fa79741d5a86