Skip to content
NuroNuroMountain View, CA

Applied AI Researcher, Agent Systems & Evaluation

Conduct research and build evaluation systems for autonomous AI agents operating on real engineering workflows. The role combines production experimentation, statistical measurement, automated optimization, and post-training open-weight vision-language models using proprietary autonomous-driving data.

Salary not listed
On-siteAI Research

About the role

Responsibilities

  • Build rigorous closed-loop evaluation systems for AI agents operating within engineering workflows.
  • Improve agent performance against real-world production data, codebases, infrastructure, and engineer requests rather than relying solely on public benchmarks.
  • Own evaluation pipelines end to end, including:
    • Collecting ground truth from production traces and human accept/reject/edit signals.
    • Building representative task suites and assessing when model-based judges are trustworthy.
    • Establishing noise floors, statistical acceptance standards, anti-gaming safeguards, and experiments for production environments.
  • Automate hill climbing across prompts, context strategies, tool sets, routing, reasoning budgets, and model selection once evaluation loops are trustworthy.
  • Post-train open-source vision-language models using supervised fine-tuning and reinforcement learning with proprietary autonomous-driving data.
  • Define data requirements and work with labeling resources to create targeted training and evaluation datasets.
  • Monitor research developments in frontier AI and translate relevant findings into experiments against production workloads.
  • Investigate test-time scaling, including reasoning budgets, sampling and search strategies, verifier-guided selection, escalation, and stopping policies.
  • Quantify real-world impact and determine whether measured results support claimed improvements.
  • Partner closely with engineering to ensure evaluation systems run safely and reliably at scale.

Requirements

  • Engineering background with strong research judgment.
  • Ability to design, implement, and ship AI research systems.
  • Fluency in frontier models and model development.
  • Expertise in evaluating foundation models and agent systems composed of models, tools, memory, retries, verifiers, and human feedback.
  • Ability to establish trustworthy measurements and interpret statistical evidence.
  • Experience working with production data and ambiguous real-world tasks.
  • Hands-on experience with supervised fine-tuning and reinforcement learning for open-weight models.

Nice-to-haves

  • Experience with vision-language models.
  • Experience with autonomous-driving or other real-world operational data.
  • Experience designing labeling programs and task-specific datasets.
  • Experience with automated optimization or hill-climbing systems.
  • Experience with test-time scaling, inference-time search, verifiers, or model routing.

Skills

artificial intelligenceLLMstransformer modelsAI AgentsModel Evaluationagent evaluationsupervised fine-tuningReinforcement Learningvision-language modelsStatistical Analysistest-time scalingmodel routingPrompt EngineeringPythonMachine Learning

Similar roles

AI Research jobs
Tulip

AI Engineer

TulipSomerville, MA

Build and deploy production LLM agents and internal AI tools, partnering with business teams to turn operational needs into secure, ROI-driven solutions. The role requires 5+ years of software experience, production AI expertise, full-stack capability, and strong cross-functional communication.

Salary not listedHybrid5+ YOEAI Research
OpenAI

Researcher, Frontier Risk Mitigations

OpenAISan Francisco, CA

Researcher developing evaluations, red-teaming pipelines, and novel mitigations for frontier AI safety risks. The role requires deep technical expertise, research engineering experience, advanced training in computer science or machine learning, and proficiency in Python or similar languages.

295k – 445k/yrOn-site4+ YOEAI Research
OpenAI

Researcher, Recursive Self-Improvement Safety

OpenAISan Francisco, CA

Conduct strategic, technically rigorous research to anticipate and mitigate loss-of-control risks from increasingly capable AI systems, including recursive self-improvement. The role combines hypothesis-driven research, rapid prototyping, safety evaluations, monitoring, and institutionalizing effective interventions.

295k – 445k/yrOn-siteAI Research
Cognition

Research, Mid-Training

CognitionSan Francisco, CA

Owns mid-training for LLMs, optimizing data mixes, synthetic data pipelines, annealing schedules, and context extension to enhance reasoning, coding, and math capabilities for AI agents. Requires deep LLM pipeline expertise, hands-on large model training, and original research contributions.

Salary not listedOn-siteAI Research
Cognition

Machine Learning Researcher

CognitionSan Francisco, CA

Conducts research to build end-to-end AI software agents capable of reasoning on real-world tasks, contributing to products like Devin and Windsurf. Requires expertise in applied AI and machine learning.

Salary not listedOn-siteAI Research