# AI Researcher

**Company:** [Improbable](https://hotfix.jobs/companies/improbable)
**Location:** Remote
**Role:** AI Research
**Skills:** Artificial Intelligence, Machine Learning, LLMs, AI Agents, Llm Evaluation, Experimental Design, Tool Use, Context Management, Python, Autonomous Workflows
**Posted:** 2026-09-09

> Conduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.

## Job Description

## Responsibilities
- Design and run experiments to evaluate agent reliability, context retention, multi-step task completion, and failure modes.
- Build and own an evaluation layer that measures whether agents complete work correctly, not merely whether their outputs sound plausible.
- Research advances in LLM agents, tool use, and reliability, and translate findings into product and engineering decisions.
- Provide clear, actionable recommendations that engineering can implement.
- Prototype promising research ideas into working product features and hand off validated solutions.
- Shape the research roadmap as the company grows.

## Requirements
- Track record of applied AI/ML research, ideally involving LLM applications, agentic systems, evaluations, or production AI reliability.
- Strong experimental design and analytical skills.
- Strong engineering fundamentals and ability to prototype experiments independently.
- Deep familiarity with LLMs, tool use, context management, and agentic-system failure modes.
- Strong written communication and ability to turn research into action.
- Pragmatic, rigorous approach to determining when findings are actionable.
- Comfort working with ambiguity and unsolved problems.
- Ability to use AI agents to accelerate research and implementation.
- Willingness to ship practical solutions.

## Nice-to-haves
- Experience with LLM evaluations, hallucination mitigation, or production AI reliability at scale.
- Experience building or evaluating agentic systems, tool use, or autonomous workflows.
- Public research or engineering work, such as papers, open-source projects, blog posts, or side projects.
- Experience with product-led research focused on shipped features.

## Similar jobs

- [Research Engineer, Takeoff Intel](https://hotfix.jobs/jobs/398824a2-65cc-4e28-aaeb-26c3b6610876) - Anthropic - San Francisco, CA - $350k – $850k/yr
- [Applied AI Research Scientist](https://hotfix.jobs/jobs/93baef6f-91a8-4c62-acaa-44c3ea48b467) - Sardine - Remote
- [Researcher, Agent Safety, Oversight and System Mitigations](https://hotfix.jobs/jobs/4544e3bb-bb96-43d2-a96c-cd364b641660) - OpenAI - San Francisco, CA - $380k – $500k/yr
- [Researcher, Agent Safety, Training and Evaluations](https://hotfix.jobs/jobs/d80336da-e453-4999-9f26-85a125b679d9) - OpenAI - San Francisco, CA - $380k – $500k/yr
- [ML Researcher - Posttraining](https://hotfix.jobs/jobs/4548e069-4358-4d58-9598-f58c621bacf0) - Krea - San Francisco, CA

**Apply:** https://hotfix.jobs/jobs/077574a5-3c6d-4634-a23c-4a909ce8aa65
**Canonical:** https://hotfix.jobs/jobs/077574a5-3c6d-4634-a23c-4a909ce8aa65