Skip to content
LovableLovableStockholm, Sweden

Researcher, Post Training

Owns Lovable’s end-to-end post-training pipeline for language models, adapting reinforcement learning and preference optimization to code-generation and agent workloads. The role combines production engineering, distributed training, evaluation, deployment, and rapid experimentation.

Salary not listed
On-siteML Engineering

About the job

Responsibilities

  • Own the full lifecycle of the post-training pipeline, from data curation and training runs through evaluation and deployment.
  • Apply and adapt reinforcement learning, preference optimization, and supervised fine-tuning methods to improve models for code generation, user-intent reasoning, and reliable agent behavior.
  • Build evaluation and experimentation infrastructure covering helpfulness, safety, latency, and reliability.
  • Develop and operate production systems for large-scale training jobs, including GPU orchestration and data pipelines.
  • Collaborate with agent, product, and infrastructure engineers to turn model gains into product improvements.
  • Investigate and resolve failures end to end, including training recipes, data issues, and serving regressions.
  • Read research papers, run experiments, and move promising research into production quickly.

Requirements

  • Hands-on experience running post-training jobs on large language models, including RFT/RLVR, preference optimization, or similar methods.
  • Strong production software engineering skills.
  • Fluency in at least one major ML framework, such as PyTorch or JAX.
  • Experience with distributed training setups and GPU clusters.
  • Understanding of the mathematics behind preference optimization, reward modeling, and alignment techniques.
  • Experience building evaluation systems that measure real-world quality beyond benchmark scores.
  • Ability to trace model-quality regressions from user-facing symptoms through serving, inference, and training.
  • Strong execution and focus on shipping improvements to users.

Nice-to-haves

  • Experience with code-generation or agentic use cases.
  • Experience putting post-trained models into the hands of real users at scale.
  • Ownership of the full loop: data curation, training, evaluation, deployment, and production monitoring.
  • Experience rapidly prototyping ideas from research papers.
  • Experience with speculative decoding or similar model-efficiency techniques.
  • Strong opinions on evaluation methodology and experience building evaluations that predict user satisfaction.
  • Meaningful contributions to the open-source ML ecosystem.

Skills

PyTorchJAXReinforcement LearningPreference OptimizationReward ModelingSupervised Fine-TuningDistributed TrainingGpu ClustersLLMsModel EvaluationData PipelinesGpu OrchestrationSpeculative DecodingCode GenerationAgentic Systems