Conduct research and build evaluation systems for autonomous AI agents operating on real engineering workflows. The role combines production experimentation, statistical measurement, automated optimization, and post-training open-weight vision-language models using proprietary autonomous-driving data.
Salary not listed
On-siteAI Research
About the role
Responsibilities
Build rigorous closed-loop evaluation systems for AI agents operating within engineering workflows.
Improve agent performance against real-world production data, codebases, infrastructure, and engineer requests rather than relying solely on public benchmarks.
Own evaluation pipelines end to end, including:
Collecting ground truth from production traces and human accept/reject/edit signals.
Building representative task suites and assessing when model-based judges are trustworthy.
Establishing noise floors, statistical acceptance standards, anti-gaming safeguards, and experiments for production environments.
Automate hill climbing across prompts, context strategies, tool sets, routing, reasoning budgets, and model selection once evaluation loops are trustworthy.
Post-train open-source vision-language models using supervised fine-tuning and reinforcement learning with proprietary autonomous-driving data.
Define data requirements and work with labeling resources to create targeted training and evaluation datasets.
Monitor research developments in frontier AI and translate relevant findings into experiments against production workloads.
Investigate test-time scaling, including reasoning budgets, sampling and search strategies, verifier-guided selection, escalation, and stopping policies.
Quantify real-world impact and determine whether measured results support claimed improvements.
Partner closely with engineering to ensure evaluation systems run safely and reliably at scale.
Requirements
Engineering background with strong research judgment.
Ability to design, implement, and ship AI research systems.
Fluency in frontier models and model development.
Expertise in evaluating foundation models and agent systems composed of models, tools, memory, retries, verifiers, and human feedback.
Ability to establish trustworthy measurements and interpret statistical evidence.
Experience working with production data and ambiguous real-world tasks.
Hands-on experience with supervised fine-tuning and reinforcement learning for open-weight models.
Nice-to-haves
Experience with vision-language models.
Experience with autonomous-driving or other real-world operational data.
Experience designing labeling programs and task-specific datasets.
Experience with automated optimization or hill-climbing systems.
Experience with test-time scaling, inference-time search, verifiers, or model routing.
Build and deploy production LLM agents and internal AI tools, partnering with business teams to turn operational needs into secure, ROI-driven solutions. The role requires 5+ years of software experience, production AI expertise, full-stack capability, and strong cross-functional communication.
Salary not listedHybrid5+ YOEAI Research
Researcher, Frontier Risk Mitigations
OpenAISan Francisco, CA
Researcher developing evaluations, red-teaming pipelines, and novel mitigations for frontier AI safety risks. The role requires deep technical expertise, research engineering experience, advanced training in computer science or machine learning, and proficiency in Python or similar languages.
295k – 445k/yrOn-site4+ YOEAI Research
Researcher, Recursive Self-Improvement Safety
OpenAISan Francisco, CA
Conduct strategic, technically rigorous research to anticipate and mitigate loss-of-control risks from increasingly capable AI systems, including recursive self-improvement. The role combines hypothesis-driven research, rapid prototyping, safety evaluations, monitoring, and institutionalizing effective interventions.
295k – 445k/yrOn-siteAI Research
Research, Mid-Training
CognitionSan Francisco, CA
Owns mid-training for LLMs, optimizing data mixes, synthetic data pipelines, annealing schedules, and context extension to enhance reasoning, coding, and math capabilities for AI agents. Requires deep LLM pipeline expertise, hands-on large model training, and original research contributions.
Salary not listedOn-siteAI Research
Machine Learning Researcher
CognitionSan Francisco, CA
Conducts research to build end-to-end AI software agents capable of reasoning on real-world tasks, contributing to products like Devin and Windsurf. Requires expertise in applied AI and machine learning.