Researcher, Post Training
Owns Lovable’s end-to-end post-training pipeline for language models, adapting reinforcement learning and preference optimization to code-generation and agent workloads. The role combines production engineering, distributed training, evaluation, deployment, and rapid experimentation.
Salary not listed
On-siteML Engineering
About the job
Responsibilities
- Own the full lifecycle of the post-training pipeline, from data curation and training runs through evaluation and deployment.
- Apply and adapt reinforcement learning, preference optimization, and supervised fine-tuning methods to improve models for code generation, user-intent reasoning, and reliable agent behavior.
- Build evaluation and experimentation infrastructure covering helpfulness, safety, latency, and reliability.
- Develop and operate production systems for large-scale training jobs, including GPU orchestration and data pipelines.
- Collaborate with agent, product, and infrastructure engineers to turn model gains into product improvements.
- Investigate and resolve failures end to end, including training recipes, data issues, and serving regressions.
- Read research papers, run experiments, and move promising research into production quickly.
Requirements
- Hands-on experience running post-training jobs on large language models, including RFT/RLVR, preference optimization, or similar methods.
- Strong production software engineering skills.
- Fluency in at least one major ML framework, such as PyTorch or JAX.
- Experience with distributed training setups and GPU clusters.
- Understanding of the mathematics behind preference optimization, reward modeling, and alignment techniques.
- Experience building evaluation systems that measure real-world quality beyond benchmark scores.
- Ability to trace model-quality regressions from user-facing symptoms through serving, inference, and training.
- Strong execution and focus on shipping improvements to users.
Nice-to-haves
- Experience with code-generation or agentic use cases.
- Experience putting post-trained models into the hands of real users at scale.
- Ownership of the full loop: data curation, training, evaluation, deployment, and production monitoring.
- Experience rapidly prototyping ideas from research papers.
- Experience with speculative decoding or similar model-efficiency techniques.
- Strong opinions on evaluation methodology and experience building evaluations that predict user satisfaction.
- Meaningful contributions to the open-source ML ecosystem.
Skills
PyTorchJAXReinforcement LearningPreference OptimizationReward ModelingSupervised Fine-TuningDistributed TrainingGpu ClustersLLMsModel EvaluationData PipelinesGpu OrchestrationSpeculative DecodingCode GenerationAgentic Systems