# Member Of Technical Staff - RL Environments

**Company:** [Cohere](https://hotfix.jobs/companies/cohere)
**Location:** London, United Kingdom
**Role:** ML Engineering
**Skills:** Reinforcement Learning, AI Agents, Synthetic Data, Model Training, Agent Evaluation, Annotation Workflows, Data Quality, Reward Design, Verifiers, Python
**Posted:** 2026-08-17

> Build and evaluate reinforcement learning environments and AI agents for enterprise workflows. The role focuses on agent training, capability measurement, verifier and reward design, synthetic data pipelines, and improving agent performance through systematic evaluation.

## Job Description

## Responsibilities
- Build new reinforcement learning (RL) environments targeting different agent capabilities and industry areas.
- Train and evaluate agents in those environments.
- Integrate tasks, data, tool implementations, and verifiers.
- Work across modeling and product to identify gaps in agent performance and improve agents and environments.
- Work with external vendors to create expert-built RL environments and tools that ensure task, data, and verifier quality.
- Automate discovery of model capability gaps and systematically measure agent performance during evaluations and training.

## Requirements
- Experience engineering agents and optimizing them for specific industry use cases.
- Experience reviewing agent trajectories to identify failure points and resolving them through model training or harness engineering.
- Strong focus on measuring agent capabilities and turning measurements into repeatable processes.
- Experience defining desirable agent outcomes, implementing verifiers, and tuning reward designs.
- Experience designing and running annotation workflows to analyze agent performance and verify data quality.
- Experience building synthetic data pipelines to scale evaluation and training efforts.
- Regular use of AI agents and experience improving agent workflows and productivity.

## Nice-to-haves
- Experience training with reinforcement learning, including scaling, troubleshooting, and tuning environments.

## Compensation and Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Full health and dental benefits, including a separate mental health budget.
- RRSP matching, 401(k), and pension scheme.
- Up to six months of 100% parental leave top-up for either parent.
- Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation (30 working days).
- Travel budget for remote employees to visit other offices and an annual company offsite.
- Coworking benefit for employees not near an office.
- $500 home office stipend.

## Similar jobs

- [AI Software Engineer](https://hotfix.jobs/jobs/c9a0e889-a36e-488f-87d0-e9c27ab63bf0) - Rollstack - Remote
- [Applied AI Engineer, Digital Natives](https://hotfix.jobs/jobs/74eb3e31-bab1-479e-90e1-2e172658343a) - OpenAI - London, United Kingdom
- [Agent Engineer](https://hotfix.jobs/jobs/627cc378-6ac6-467e-bbf0-2c07738ae0a5) - Elliptic - London, United Kingdom
- [AI Engineer - New Verticals](https://hotfix.jobs/jobs/4159c72e-536e-4211-969c-6bcb9ad805fd) - Protege - Remote
- [AI Engineer - Assistant Experience](https://hotfix.jobs/jobs/00d11204-96cf-4c85-a12f-9058f0581c15) - Build - New York, NY - $120k – $240k/yr

**Apply:** https://hotfix.jobs/jobs/def8fba7-a52f-4ec2-b89b-0ba0974ea9bc
**Canonical:** https://hotfix.jobs/jobs/def8fba7-a52f-4ec2-b89b-0ba0974ea9bc