Research Engineer, Reinforcement Learning
Develops RL environments and fine-tunes language models using PPO, DPO, and KTO to enhance agentic capabilities for data infrastructure tasks. Requires deep RL expertise, LLM fine-tuning knowledge, and strong problem-solving skills.
About the job
Responsibilities
- Develop and refine reward functions to optimize agent behavior for complex data engineering tasks.
- Create RL gym environments for language model agents.
- Fine-tune language models using reinforcement learning techniques such as PPO, DPO, and KTO.
- Stay at the forefront of research on RL for language models, incorporating advancements like GRPO, SWE-Gym, and SWE-RL into practical applications.
- Curate and build high-quality datasets for supervised fine-tuning (SFT) and RLHF.
- Design experiments to evaluate and improve the agentic capabilities of language models in data environments.
Requirements
- Deep understanding of reinforcement learning, reward shaping, and optimization strategies.
- Strong familiarity with LLM fine-tuning techniques (PPO, DPO, KTO) and their applications in reinforcement learning.
- Knowledge of recent advancements in RL for language models (GRPO, SWE-Gym, SWE-RL).
- Experience curating and constructing high-quality datasets for fine-tuning.
- Strong problem-solving skills and a history of working on complex ML projects.
- High agency—ability to work independently, experiment proactively, and drive research initiatives forward.
Nice-to-Haves
- Experience with distributed training in PyTorch (DDP, FSDP).
- Hands-on experience designing RL environments for traditional RL problems.
- Contributions to open-source projects in RL, LLMs, or ML infrastructure.
- Familiarity with data lakes and warehouses (Snowflake, BigQuery, Redshift).
Benefits
- 100% employer-covered health, dental, and vision insurance.
- 401(k) with company match.
- Access to Bay Club or Equinox in San Francisco.
Skills
Reinforcement Learning, Ppo, Dpo, Kto, RLHF, PyTorch, Gym, Grpo, Swe-Gym, Swe-Rl
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.