Research Engineer, Computer Use
Research Engineer advancing Claude's computer use capabilities through experiments, RL environments, evaluations, and infrastructure for perception and agentic tasks. Requires Python, ML training/evaluation experience, and a focus on safe AI.
About the job
Key Responsibilities
- Design and run experiments to improve Claude's perception and agentic capabilities
- Develop robust, reliable evaluation frameworks for measuring our models' ability to complete complex computer tasks
- Build and improve computer use and vision reinforcement learning training environments
- Create pipelines and tools to test and validate complex RL environments
- Collaborate with teams across the model training and infrastructure stack to improve our production training setup
- Partner with product teams to bring research advances into production
Minimum Qualifications
- Software engineering experience and proficiency in Python
- Experience training, fine-tuning, or evaluating machine learning models
- Strong communication skills and a collaborative working style
- Care about the societal impacts and safety of your work
Preferred Qualifications
- Experience training models for computer use or other agentic capabilities
- Experience with reinforcement learning, particularly in long-horizon or sparse-reward settings
- Familiarity with multimodal model training
- Experience building evaluations or benchmarks for agentic systems
- Experience building reinforcement learning environments, simulation systems, or large-scale ML infrastructure
- Experience working closely with product teams to drive model improvements
Education
- Bachelor’s degree or an equivalent combination of education, training, and/or experience in a field relevant to the role
Skills
Python, Machine Learning, Reinforcement Learning, Multimodal Models, Rl Environments, Model Training, Model Evaluation, Model Fine-Tuning, Agentic Systems, Benchmarks, ML Infrastructure
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Operates and improves the infrastructure powering large-scale post-training and reinforcement learning runs, partnering with researchers to debug failures, improve reliability, and automate recovery. Requires 4+ years operating distributed production systems and strong Python, Go, or C++ skills.
Research-focused engineer advancing agentic model capabilities across synthetic data, task environments, evaluations, training, and usability improvements. Requires strong Python engineering, deep learning framework experience, scalable distributed training skills, and scientific experimentation ability.
Researcher focused on scaling reinforcement learning for frontier models, with ownership spanning asynchronous RL algorithms, inference and distributed training systems, and large-scale empirical studies. Requires strong Python and deep learning experience, scalable systems debugging, and rigorous research judgment.
Develop multimodal perception and authentication systems combining visual, audio, and other sensor signals for real-world AI products. The role requires machine learning expertise, practical research-to-system experience, and proficiency in Python and PyTorch with comfort in C++.