Research Scientist
Research Scientist defining and executing research on reliable long-horizon agents in enterprise environments. The role focuses on post-training and reinforcement learning, agent memory, evaluation, verification, and structured representations, combining hands-on experimentation with product delivery and publication.
About the job
Responsibilities
- Set and pursue a research agenda for reliable, long-horizon agentic behavior in enterprise environments.
- Invent and validate post-training methods for transformation work, including reward design, reinforcement learning, long-horizon planning, tool use, curriculum, and data strategy.
- Design memory architectures for multi-step, multi-day agent runs, including persistence, retrieval, revision, and model training.
- Develop evaluation methodologies that predict customer-observed correctness and identify failures in automated proxies.
- Study multi-agent failure modes, including error compounding, partial observability, delegation, and verification.
- Research verification methods for changes where automated tests do not provide coverage.
- Develop enterprise representations such as ontologies and knowledge graphs for reliable agent reasoning.
- Investigate post-training compute and data scaling across affordable open-weight models.
- Turn research findings into shipped product capabilities with Research Engineering and product teams.
- Publish papers, technical reports, and open-source artifacts; represent the company’s research externally.
- Review experiment designs and mentor engineers moving into research.
Requirements
- MS or PhD in computer science, machine learning, statistics, mathematics, physics, or a related field, or comparable research experience.
- Track record of original machine learning research through publications, influential open-source work, or impactful lab results.
- Deep expertise in at least one of post-training and reinforcement learning for LLMs, agent memory and long-context reasoning, knowledge representation and structured reasoning, or evaluation methodology.
- Hands-on ability to write code, run experiments, and analyze logs.
- Exceptional research judgment, rigor, urgency, and technical communication.
Nice-to-haves
- Experience owning a research direction at a frontier lab or strong academic group.
- Published work on agents, tool use, reinforcement learning for LLMs, code generation or repair, reasoning, or evaluation.
- Experience with verifiable-reward reinforcement learning or incomplete-verification reward-model failure modes.
- Experience with knowledge representation, ontologies, or neurosymbolic methods.
Compensation
- Annual salary: $200,000–$300,000.
Skills
Machine Learning, Reinforcement Learning, LLMs, Post-Training, Agent Memory, Long-Context Reasoning, Knowledge Representation, Ontologies, Knowledge Graphs, Evaluation Methodology, Multi-Agent Systems, Python, Open-Weight Models, Code Generation
Similar jobs
AI Research jobsLeads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.
Conduct rigorous people research and applied data science to evaluate talent programs, organizational health, and employee experiences. The role requires advanced expertise in research design, experimentation, measurement, causal inference, statistical modeling, and responsible handling of sensitive employee data.
Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.
Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.
Research novel post-training methods for large language models, focusing on preference optimization, data curation, evaluation, alignment, and robustness across text and multimodal systems. Requires advanced academic training and experience with deep learning, reinforcement learning, and post-training techniques.