Evals Infrastructure Tech Lead / Manager
Lead the Evals Infrastructure team at Anthropic building large-scale distributed systems for model evaluation, orchestration, and trustworthy metrics that inform launch decisions for frontier AI models. Requires strong distributed systems experience, Python/Rust proficiency, and people management skills with a focus on measurement quality and AI safety.
About the job
Responsibilities
- Lead the team building the distributed systems that schedule, orchestrate, and execute evals for our frontier model training
- Own eval throughput and cost: compute allocation across suites, queueing against constrained accelerator pools, caching and reuse of eval work
- Build and scale the harnesses researchers use to define, run, and iterate on evals
- Make eval results trustworthy — determinism, reproducibility, and honest uncertainty quantification on reported metrics
- Ensure eval signal reaches the dashboards and reviews where launch decisions actually get made
- Contribute directly as an engineer while managing and growing the team, prioritizing its work, and coaching your reports
Requirements
- Led technical projects end-to-end on large-scale distributed systems, and have 1+ years managing engineers (or tech-lead-with-reports experience)
- Strong in Python and Rust
- Built high-throughput, fault-tolerant systems on cloud or on-prem accelerator fleets
- Care about measurement quality, not just pipeline uptime — you'd notice if a metric moved for the wrong reason
- Communicate well with researchers and can translate research needs into infrastructure
- Deeply interested in the transformative effects of advanced AI and committed to safe development
- Bachelor’s degree or an equivalent combination of education, training, and/or experience in a field relevant to the role
Nice-to-Haves
- Worked on LLM inference or training infrastructure
- Experience with eval or benchmarking systems, especially agentic evals requiring sandboxed execution
- Working statistical literacy — variance, confidence intervals, sample-size sufficiency for noisy metrics
- Experience with observability and regression detection over time-series metrics
Skills
Python, Rust, Distributed Systems, Llm Inference, Llm Training, Observability, Sandboxing
Similar jobs
Engineering Management jobsLeads a research engineering team developing biological safety evaluations, datasets, and classifiers for frontier AI models. The role combines technical direction and people management with deep life-sciences, machine-learning, experimentation, and biosecurity expertise.
Leads the team responsible for Anthropic’s compute scheduling platform, job-launch tooling, and fleet-efficiency systems. The role requires engineering management experience, a hands-on software engineering background, and expertise in large-scale infrastructure and Kubernetes scheduling.
Leads Anthropic’s Network Security Engineering team, setting strategy and managing engineers who build secure-by-default controls across cloud providers and AI data centers. The role requires extensive engineering management experience, hands-on network security expertise, Kubernetes networking knowledge, and a bachelor’s degree or equivalent experience.
Leads and develops an applied AI engineering team delivering reliable, customer-facing agents, integrations, and scientific infrastructure for pharma and biotech organizations. The role combines hands-on engineering, partner ownership, cross-functional product influence, and responsible AI deployment in life sciences.
Leads an engineering team building AI-native internal productivity tools, partnering with business stakeholders and owning reliable cloud services. Requires 6+ years of engineering experience, 2+ years managing engineering teams, and experience with web applications, APIs, and cloud infrastructure.