# Researcher, Post Training

**Company:** [Lovable](https://hotfix.jobs/companies/lovable)
**Location:** Stockholm, Sweden, London, United Kingdom
**Role:** ML Engineering
**Skills:** PyTorch, JAX, Reinforcement Learning, Preference Optimization, Reward Modeling, Supervised Fine-Tuning, Distributed Training, Gpu Clusters, LLMs, Model Evaluation, Data Pipelines, Gpu Orchestration, Speculative Decoding, Code Generation, Agentic Systems
**Posted:** 2026-04-28

> Owns Lovable’s end-to-end post-training pipeline for language models, adapting reinforcement learning and preference optimization to code-generation and agent workloads. The role combines production engineering, distributed training, evaluation, deployment, and rapid experimentation.

## Job Description

## Responsibilities
- Own the full lifecycle of the post-training pipeline, from data curation and training runs through evaluation and deployment.
- Apply and adapt reinforcement learning, preference optimization, and supervised fine-tuning methods to improve models for code generation, user-intent reasoning, and reliable agent behavior.
- Build evaluation and experimentation infrastructure covering helpfulness, safety, latency, and reliability.
- Develop and operate production systems for large-scale training jobs, including GPU orchestration and data pipelines.
- Collaborate with agent, product, and infrastructure engineers to turn model gains into product improvements.
- Investigate and resolve failures end to end, including training recipes, data issues, and serving regressions.
- Read research papers, run experiments, and move promising research into production quickly.

## Requirements
- Hands-on experience running post-training jobs on large language models, including RFT/RLVR, preference optimization, or similar methods.
- Strong production software engineering skills.
- Fluency in at least one major ML framework, such as PyTorch or JAX.
- Experience with distributed training setups and GPU clusters.
- Understanding of the mathematics behind preference optimization, reward modeling, and alignment techniques.
- Experience building evaluation systems that measure real-world quality beyond benchmark scores.
- Ability to trace model-quality regressions from user-facing symptoms through serving, inference, and training.
- Strong execution and focus on shipping improvements to users.

## Nice-to-haves
- Experience with code-generation or agentic use cases.
- Experience putting post-trained models into the hands of real users at scale.
- Ownership of the full loop: data curation, training, evaluation, deployment, and production monitoring.
- Experience rapidly prototyping ideas from research papers.
- Experience with speculative decoding or similar model-efficiency techniques.
- Strong opinions on evaluation methodology and experience building evaluations that predict user satisfaction.
- Meaningful contributions to the open-source ML ecosystem.

## Similar jobs

- [AI Software Engineer](https://hotfix.jobs/jobs/c9a0e889-a36e-488f-87d0-e9c27ab63bf0) - Rollstack - Remote
- [Applied AI Engineer, Digital Natives](https://hotfix.jobs/jobs/74eb3e31-bab1-479e-90e1-2e172658343a) - OpenAI - London, United Kingdom
- [Agent Engineer](https://hotfix.jobs/jobs/627cc378-6ac6-467e-bbf0-2c07738ae0a5) - Elliptic - London, United Kingdom
- [AI Engineer - New Verticals](https://hotfix.jobs/jobs/4159c72e-536e-4211-969c-6bcb9ad805fd) - Protege - Remote
- [AI Engineer - Assistant Experience](https://hotfix.jobs/jobs/00d11204-96cf-4c85-a12f-9058f0581c15) - Build - New York, NY - $120k – $240k/yr

**Apply:** https://hotfix.jobs/jobs/08746eda-0ffd-4697-ae82-08f01084eff4
**Canonical:** https://hotfix.jobs/jobs/08746eda-0ffd-4697-ae82-08f01084eff4