AI Researcher
Conduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.
About the job
Responsibilities
- Design and run experiments to evaluate agent reliability, context retention, multi-step task completion, and failure modes.
- Build and own an evaluation layer that measures whether agents complete work correctly, not merely whether their outputs sound plausible.
- Research advances in LLM agents, tool use, and reliability, and translate findings into product and engineering decisions.
- Provide clear, actionable recommendations that engineering can implement.
- Prototype promising research ideas into working product features and hand off validated solutions.
- Shape the research roadmap as the company grows.
Requirements
- Track record of applied AI/ML research, ideally involving LLM applications, agentic systems, evaluations, or production AI reliability.
- Strong experimental design and analytical skills.
- Strong engineering fundamentals and ability to prototype experiments independently.
- Deep familiarity with LLMs, tool use, context management, and agentic-system failure modes.
- Strong written communication and ability to turn research into action.
- Pragmatic, rigorous approach to determining when findings are actionable.
- Comfort working with ambiguity and unsolved problems.
- Ability to use AI agents to accelerate research and implementation.
- Willingness to ship practical solutions.
Nice-to-haves
- Experience with LLM evaluations, hallucination mitigation, or production AI reliability at scale.
- Experience building or evaluating agentic systems, tool use, or autonomous workflows.
- Public research or engineering work, such as papers, open-source projects, blog posts, or side projects.
- Experience with product-led research focused on shipped features.
Skills
Artificial Intelligence, Machine Learning, LLMs, AI Agents, Llm Evaluation, Experimental Design, Tool Use, Context Management, Python, Autonomous Workflows
Similar jobs
AI Research jobsResearch Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.
Conduct applied research on foundation models for fraud detection using large-scale behavioral and financial-risk data. The role spans experimentation, evaluation, production deployment, and cross-functional work on model governance, requiring 4+ years of applied ML experience and strong Python and SQL skills.
Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.
Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.
Conduct research and engineering on large-scale post-training of diffusion and language models, focusing on aesthetics, preference optimization, reinforcement learning, reward modeling, and evaluation. The role requires strong PyTorch, distributed training, low-precision inference, and VLM experience.