Researcher, Synthetic RL
Develops novel reinforcement learning techniques using synthetic environments and feedback to enhance large-scale AI models. Designs experiments, analyzes dynamics, and integrates research into production systems; requires strong RL/ML background and engineering skills.
About the job
In this role, you will:
- Research and develop reinforcement learning algorithms
- Design and run experiments to study training dynamics and model behavior at scale
- Collaborate with engineers and researchers to integrate successful approaches into model training pipelines
You might thrive in this role if you:
- Have a strong background in reinforcement learning, machine learning research, or related fields
- Have strong engineering and statistical analysis skills
- Enjoy exploring new problem spaces where data, objectives, and evaluation are imperfect or evolving
- Are motivated by seeing research ideas influence real-world AI systems
Skills
Reinforcement Learning, Machine Learning, Python, Statistical Analysis, Experiment Design, Synthetic Data, Self-Play, Simulators, Ai Training Pipelines, Research
Similar jobs
AI Research jobsConducts frontier AI research for health, developing and evaluating scalable training methods, models, and agents that improve medical reasoning, reliability, and real-world outcomes. Requires exceptional machine learning or biomedical AI research depth, hands-on coding and experimentation, and end-to-end ownership of ambiguous problems.
Conducts hands-on medicinal chemistry research to evaluate AI-generated molecules and synthetic routes, advancing small-molecule programs from design through experimental validation. Requires a chemistry PhD, sustained synthetic experience, and cross-functional collaboration skills.
Applied research scientists develop deep-learning and generative media systems for video, audio, and multimodal editing features that ship to millions of users. The role requires strong PyTorch or TensorFlow skills, rapid experimentation, and evidence of impactful research or production machine-learning work.
Conducts causal inference research for financial market prediction and portfolio optimization, developing and validating models from research through live trading. Requires Ph.D.-level coursework, strong causal inference and statistics expertise, mathematical ability, and production Python skills.
Research Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.