# Applied AI Researcher, Agent Systems & Evaluation

**Company:** [Nuro](https://hotfix.jobs/companies/nuro)
**Location:** Mountain View, CA, California
**Role:** AI Research
**Skills:** artificial intelligence, LLMs, transformer models, AI Agents, Model Evaluation, agent evaluation, supervised fine-tuning, Reinforcement Learning, vision-language models, Statistical Analysis, test-time scaling, model routing, Prompt Engineering, Python, Machine Learning
**Posted:** 2026-08-11

> Conduct research and build evaluation systems for autonomous AI agents operating on real engineering workflows. The role combines production experimentation, statistical measurement, automated optimization, and post-training open-weight vision-language models using proprietary autonomous-driving data.

## Job Description

## Responsibilities

- Build rigorous closed-loop evaluation systems for AI agents operating within engineering workflows.
- Improve agent performance against real-world production data, codebases, infrastructure, and engineer requests rather than relying solely on public benchmarks.
- Own evaluation pipelines end to end, including:
  - Collecting ground truth from production traces and human accept/reject/edit signals.
  - Building representative task suites and assessing when model-based judges are trustworthy.
  - Establishing noise floors, statistical acceptance standards, anti-gaming safeguards, and experiments for production environments.
- Automate hill climbing across prompts, context strategies, tool sets, routing, reasoning budgets, and model selection once evaluation loops are trustworthy.
- Post-train open-source vision-language models using supervised fine-tuning and reinforcement learning with proprietary autonomous-driving data.
- Define data requirements and work with labeling resources to create targeted training and evaluation datasets.
- Monitor research developments in frontier AI and translate relevant findings into experiments against production workloads.
- Investigate test-time scaling, including reasoning budgets, sampling and search strategies, verifier-guided selection, escalation, and stopping policies.
- Quantify real-world impact and determine whether measured results support claimed improvements.
- Partner closely with engineering to ensure evaluation systems run safely and reliably at scale.

## Requirements

- Engineering background with strong research judgment.
- Ability to design, implement, and ship AI research systems.
- Fluency in frontier models and model development.
- Expertise in evaluating foundation models and agent systems composed of models, tools, memory, retries, verifiers, and human feedback.
- Ability to establish trustworthy measurements and interpret statistical evidence.
- Experience working with production data and ambiguous real-world tasks.
- Hands-on experience with supervised fine-tuning and reinforcement learning for open-weight models.

## Nice-to-haves

- Experience with vision-language models.
- Experience with autonomous-driving or other real-world operational data.
- Experience designing labeling programs and task-specific datasets.
- Experience with automated optimization or hill-climbing systems.
- Experience with test-time scaling, inference-time search, verifiers, or model routing.

## Similar roles

- [AI Engineer](https://hotfix.jobs/jobs/db221416-03ed-4e4b-a95e-06b4bd4dc66f) - Tulip - Somerville, MA
- [Researcher, Frontier Risk Mitigations](https://hotfix.jobs/jobs/7ff5dad3-4092-4016-af83-7697cc0a93a8) - OpenAI - San Francisco, CA - $295k – $445k/yr
- [Researcher, Recursive Self-Improvement Safety](https://hotfix.jobs/jobs/1f8fd130-4760-4d3a-af28-47681ec30870) - OpenAI - San Francisco, CA - $295k – $445k/yr
- [Research, Mid-Training](https://hotfix.jobs/jobs/ed66eec5-bae2-473e-a2eb-f85b9ea6a95c) - Cognition - San Francisco, CA
- [Machine Learning Researcher](https://hotfix.jobs/jobs/ff556b68-ccd8-4ad8-b9ac-00201cc76132) - Cognition - San Francisco, CA

**Apply:** https://hotfix.jobs/jobs/baef7949-4d56-4db9-9fb2-9426715cd248
**Canonical:** https://hotfix.jobs/jobs/baef7949-4d56-4db9-9fb2-9426715cd248