Research Engineer, LangSmith Engine
Research Engineer improving the capability, efficiency, and reliability of autonomous AI agents through benchmarks, experiments, prompting, model optimization, and post-training. The role requires 4+ years of ML/AI research experience, strong software engineering skills, and a master’s or Ph.D. in a relevant scientific field.
About the job
Responsibilities
- Build and maintain benchmarks and evaluations that measure the quality and efficiency of Engine agents on real-world tasks.
- Design and run experiments to improve agent performance across models, prompting, context, tools, orchestration, and agent strategies.
- Explore and implement post-training and fine-tuning techniques when they can meaningfully improve agent capabilities, quality, or cost.
- Turn successful experiments into production improvements, working closely with engineers and researchers to measure impact and prevent regressions.
- Help define the ML roadmap and technical direction for improving Engine agents.
- Mentor other engineers through strong technical leadership.
Requirements
- 4+ years of experience in ML/AI research or a closely related field.
- Master’s or Ph.D. in a relevant scientific field.
- Hands-on experience working with LLMs and AI agents, including analyzing model behavior and improving real-world performance.
- Strong experience designing benchmarks, evaluations, and experiments for AI/ML systems.
- Strong software engineering skills, with a track record of taking ideas from research prototype to measurable production impact.
- Strong research judgment, ability to work through ambiguity, move quickly, and communicate findings clearly.
Nice to Have
- Ph.D. in Machine Learning, Computer Science, or Physics.
- Experience with LLM-as-a-judge, automated graders, synthetic data generation, or human evaluation.
- Experience with reinforcement learning, preference optimization, SFT, RLHF/RLAIF, or other post-training techniques for LLMs.
- Experience optimizing LLM agents for cost, latency, or task efficiency in production.
- Experience with model serving, inference optimization, distributed systems, or GPU infrastructure.
Compensation and Benefits
- Competitive compensation including base salary, variable compensation for relevant roles, meaningful equity, benefits, and perks.
- Medical, dental, and vision coverage.
- Flexible vacation.
- 401(k) plan.
- Life insurance.
- Meals on in-office days in the US.
- Benefits and offerings vary by role, level, and location.
Skills
Machine Learning, Artificial Intelligence, LLMs, AI Agents, Benchmarking, Model Evaluation, Prompt Engineering, Fine-Tuning, Reinforcement Learning, Preference Optimization, Sft, RLHF, Model Serving, Distributed Systems, Gpu Infrastructure
Similar jobs
AI Research jobsConduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.
Research Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.
Conduct applied research on foundation models for fraud detection using large-scale behavioral and financial-risk data. The role spans experimentation, evaluation, production deployment, and cross-functional work on model governance, requiring 4+ years of applied ML experience and strong Python and SQL skills.
Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.
Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.