Member of Technical Staff, Enterprise Evals Platform
Builds the platform, verifiers, environments, and grading infrastructure used to evaluate enterprise AI agents at scale. The role combines strong software engineering with expertise in agent runtimes, evaluation design, benchmarks, and production failure analysis.
About the job
Responsibilities
- Define golden sets by decomposing real tasks and encoding expert quality standards.
- Build verifiers over agent trajectories and outputs that are calibrated and difficult to game.
- Build evaluation platforms that run offline environments, task suites, and grading infrastructure at scale.
- Run loss analysis over production trajectories and turn failure modes into regression tests.
- Run optimization loops across models, prompts, skills, and harnesses.
- Own rollout gates that determine whether an agent change ships.
- Partner with the Enterprise Platform team and Applied AI engineers embedded with customers.
Requirements
- Professional, academic, or research experience in agent engineering and evaluation, including agent runtimes, harnesses, trajectories, and failure modes.
- Experience building evaluation suites for LLM or agent systems.
- Familiarity with benchmarks such as terminal-bench, tau-bench, and APEX, including how they are constructed and where they can be gamed.
- Strong judgment about task and rubric design, translating fuzzy notions of quality into measurable criteria.
- Strong software engineering fundamentals.
- Ability to work independently on ambiguous, loosely specified problems.
Nice-to-haves
- Experience with Harbor environments and reinforcement learning environments.
Compensation and Benefits
- Salary range: $220,000–$425,000.
- Up to $15,000 relocation bonus.
- $10,000 housing bonus for employees living within 0.5 miles of the office.
- $1,500 monthly meal stipend.
- Generous equity grant vested over four years.
- Free Equinox membership.
- $200 monthly laundry reimbursement.
- $200 monthly personal wellness reimbursement.
- Health, dental, and vision insurance.
Skills
Agent Engineering, Llm Evaluation, Evaluation Suites, Software Engineering, Agent Runtimes, Agent Harnesses, Benchmarking, Terminal-Bench, Tau-Bench, Apex, Harbor, Reinforcement Learning, Regression Testing, Prompt Optimization, Python
Similar jobs
ML Engineering jobsBuild and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.
Build AI agent harnesses, models, and product capabilities that enable agents to perform complex work across digital environments. The role combines applied AI research and software engineering, requiring Python proficiency, strong product judgment, and experience with agent tooling, reinforcement learning, or browser technologies.
Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.
Build and operate production machine-learning systems for search ranking, relevance, extraction quality, and LLM-driven features. The role requires production ML ownership, ranking or relevance expertise, large-scale data experience, Python, and rigorous experimentation skills.
Build and deploy algorithmic systems for high-impact healthcare problems, choosing among machine learning, optimization, heuristics, and hybrid approaches. The role requires 4+ years of relevant industry experience, strong applied problem-solving and evaluation skills, and fluency in modern ML tooling.