Lead the Evals Infrastructure team at Anthropic building large-scale distributed systems for model evaluation, orchestration, and trustworthy metrics that inform launch decisions for frontier AI models. Requires strong distributed systems experience, Python/Rust proficiency, and people management skills with a focus on measurement quality and AI safety.
500k – 850k/yr
Hybrid7+ YOEEngineering Management
About the role
Responsibilities
Lead the team building the distributed systems that schedule, orchestrate, and execute evals for our frontier model training
Own eval throughput and cost: compute allocation across suites, queueing against constrained accelerator pools, caching and reuse of eval work
Build and scale the harnesses researchers use to define, run, and iterate on evals
Make eval results trustworthy — determinism, reproducibility, and honest uncertainty quantification on reported metrics
Ensure eval signal reaches the dashboards and reviews where launch decisions actually get made
Contribute directly as an engineer while managing and growing the team, prioritizing its work, and coaching your reports
Requirements
Led technical projects end-to-end on large-scale distributed systems, and have 1+ years managing engineers (or tech-lead-with-reports experience)
Strong in Python and Rust
Built high-throughput, fault-tolerant systems on cloud or on-prem accelerator fleets
Care about measurement quality, not just pipeline uptime — you'd notice if a metric moved for the wrong reason
Communicate well with researchers and can translate research needs into infrastructure
Deeply interested in the transformative effects of advanced AI and committed to safe development
Bachelor’s degree or an equivalent combination of education, training, and/or experience in a field relevant to the role
Nice-to-Haves
Worked on LLM inference or training infrastructure
Experience with eval or benchmarking systems, especially agentic evals requiring sandboxed execution
Working statistical literacy — variance, confidence intervals, sample-size sufficiency for noisy metrics
Experience with observability and regression detection over time-series metrics
Lead the product engineering organization for Statsig at OpenAI, defining strategy for experimentation, feature rollout, configuration, and analytics platforms. Build and scale leadership teams while partnering with product, research, and infrastructure groups to turn launch and measurement needs into reliable company-wide capabilities.
441k – 490k/yr
Hybrid8+ YOEEngineering Management
Engineering Manager, Research Data Platform
AnthropicSan Francisco, CA +1
Lead the Research Data Platform team at Anthropic as technical lead. Set direction for data systems and canonical datasets that researchers rely on, own end-to-end pipelines and platform components, and drive adoption through close collaboration with research teams. Requires experience building scalable data platforms and setting technical direction.
405k – 850k/yr
Hybrid7+ YOEEngineering Management
Engineering Manager, Cybersecurity Products
AnthropicSan Francisco, CA +1
Lead an engineering team building AI-powered cybersecurity products, focusing on prototyping, shipping, and scaling solutions. This role involves technical leadership, customer engagement, and architectural decisions across the full stack.
405k – 485k/yr
Hybrid8+ YOEEngineering Management
Engineering Manager
Thinking Machines LabSan Francisco, CA
Leads a team of senior/staff engineers building scalable ML infrastructure and products, owning system design, reliability, and execution while contributing hands-on and hiring top talent. Requires 8+ years in production systems and 3+ years managing engineers.
400k – 500k/yr
On-site8+ YOEEngineering Management
Senior Enterprise Sales Manager | Housing
EliseAINew York, NY +1
Leads enterprise sales team for housing SaaS platform, driving new business and expansions with property management companies. Requires 4+ years sales management experience in enterprise B2B SaaS and onsite office presence.