Staff Research Engineer, Discovery Team
Research Engineer works end-to-end to remove bottlenecks toward scientific AGI, focusing on long-horizon reasoning, computer use, and model capabilities. Requires 8+ years ML experience, expertise in language model pipelines, distributed systems, and collaborative problem-solving.
About the job
About the Team
Our team is focused on building an AI scientist capable of solving long-term reasoning challenges and pushing the scientific frontier. We're currently improving models' abilities to use computers as a laboratory for long-horizon tasks and scientific workflows.
About the Role
As a Research Engineer, you will work end-to-end to identify and address key blockers on the path to scientific AGI. Strong candidates have familiarity with language model training, evaluation, and inference, and enjoy collaborative problem-solving.
Responsibilities
- Working across the full stack to identify and remove bottlenecks preventing progress toward scientific AGI
- Develop approaches to address long-horizon task completion and complex reasoning challenges essential for scientific discovery
- Scaling research ideas from prototype to production
- Create benchmarks and evaluation frameworks to measure model capabilities in scientific workflows and computer use
- Implement distributed training systems and performance optimizations to support large-scale model development
You may be a good fit if you
- Have 8+ years of ML research experience
- Are familiar with large scale language model training, evaluation, and inference pipelines
- Enjoy obsessively iterating on immediate blockers towards long-term goals
- Thrive working collaboratively to solve problems
- Have expertise in performance optimization and distributed computing systems
- Show strong problem-solving skills and ability to identify technical bottlenecks in complex systems
- Can translate research concepts into scalable engineering solutions
- Have a track record of shipping ML systems that tackle challenging multi-step reasoning problems
Strong candidates may also have
- Expertise with performance optimization for language model inference and training
- Experience with computer use automation and agentic AI systems
- A history working on reinforcement learning approaches for complex task completion
- Knowledge of containerization technologies (Docker, Kubernetes) and cloud deployment at scale
- Demonstrated ability to work across multiple domains (language modeling, systems engineering, scientific computing)
- Experience with VM/sandboxing/container deployment and large-scale data processing
- Experience working with large scale data problem solving and infrastructure
- Published research or practical experience in scientific AI applications or long-horizon reasoning
Logistics
Education requirements: At least a Bachelor's degree in a related field or equivalent experience.
Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. Some roles may require more time in our offices.
Skills
PyTorch, Transformers, Distributed Training, Kubernetes, Docker, Performance Optimization, Reinforcement Learning, Language Models, Inference Optimization, Data Pipelines
Similar jobs
AI Research jobsResearch and evaluate frontier AI capabilities for cybersecurity, rapidly prototyping tools, designing rigorous benchmarks, and helping operationalize reliable capabilities into products. Requires deep security expertise, strong technical communication, and at least seven years of relevant experience.
Leads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.
Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.
Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.
Own the architecture, delivery, evaluation, and production operations of AI capabilities embedded in procurement and finance workflows. The role requires 10+ years in applied AI or machine learning, deep LLM and agent expertise, and experience delivering measurable production outcomes.