Core Responsibilities
Frontier Model Evaluation & Data Strategy
- Develop a working understanding of what our customers' models can and can't do, and where the capability gaps are.
- Evaluate model outputs and benchmark performance, and create targeted loss analysis to identify what data will drive improvement.
- Translate research goals into concrete task designs, difficulty targets, and quality specifications — including novel frontier tasks that don't exist anywhere else.
Pipeline Design & Innovation
- Design how data is produced — human-expert, synthetic, and hybrid model-in-the-loop pipelines — and continuously improve them.
- Prototype new generation and validation approaches; bring ideas for new data products, not just improvements to existing ones.
- Balance quality, throughput, and cost as you scale a pipeline from prototype to production.
Operational Excellence
- Manage end-to-end data pipelines from customer specification to final delivery.
- Diagnose bottlenecks, restructure workflows, and implement solutions — incentive systems, workflow re-sequencing, sharper instructions, scaled review processes, and automated quality assurance.
- Run daily internal syncs ("war rooms") to stay ahead of issues.
Customer Relationships
- Act as the primary point of contact for leading AI labs; deliver clear, consistent reporting.
- Proactively anticipate researcher needs and identify opportunities for expansion.
Expert Teams & Workflow Design
- Design the human-expert labeling and evaluation flows on the platform — how tasks are structured, reviewed, and scored.
- Source, vet, train, and performance-manage teams of domain experts (software engineers, competitive programmers, and specialists).
- Maintain a high execution and quality standard across every stage of production.
Requirements
- Technical judgment: Coding literacy and genuine ML or Model Benchmark familiarity — enough to evaluate model outputs, reason about frontier-model capabilities and failure modes, read benchmark/eval work, and translate research goals into concrete task and data designs.
- Operational ownership: A track record of running complex, high-stakes projects end to end, and genuine energy for large-scale execution and gritty process optimization under pressure.
- Communication & customer instinct: Strong analytical and communication skills; comfortable owning high-profile relationships with technical customers.
Nice-to-Haves
- Pipeline building: designing data or automation pipelines; hands-on with LLMs/agents; building synthetic-data or model-in-the-loop systems.
- Research fluency: connecting model/benchmark literature to what data would move a frontier model; having created a benchmark or published analysis of model behavior.
Backgrounds that often fit: ML/data/software engineers who love operating, technical PMs, research engineers, or strong generalist operators with real technical range. Backgrounds from consulting, finance, or high-growth startups can work if paired with real technical fluency.
Compensation & Benefits
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance