# Research Engineer

**Company:** [LiveKit](https://hotfix.jobs/companies/livekit)
**Location:** Remote
**Role:** ML Engineering
**Salary:** $135k – $300k/yr
**Skills:** Python, Reinforcement Learning, Grpo, Fine-Tuning, Trl, Verl, Openrlhf, vLLM, Sglang, Fsdp, Qwen, Llama, Lora, Gpu Computing, Multi-Turn Agents
**Posted:** 2026-08-18

> Build and productionize post-training systems for voice and text agents, including environments, verifiers, synthetic data, evaluations, and model training. The role requires strong Python and end-to-end model development experience, with reinforcement learning and distributed training expertise preferred.

## Job Description

## Responsibilities
- Build the environments and verifiers used to train models.
- Own the synthetic data pipeline from generation through quality gates.
- Run end-to-end training experiments and explain model changes.
- Build release evaluations and define the criteria models must meet.
- Choose and adapt open-weight base models for assigned tasks.
- Ensure trained behavior works across voice and text agents.
- Ship models to production and improve them using real-world usage.

## Requirements
- Strong Python engineering skills.
- Experience carrying a model from raw data through production.
- Strong data quality practices, including coverage, diversity, and leakage prevention.
- Ability to design against weak or exploitable rewards.
- Comfort working with GPUs and understanding their limitations.
- Ability to determine when to train and when not to.
- Comfortable collaborating in a remote environment.

## Nice to Have
- Experience with post-training, including fine-tuning, reward design, or reinforcement learning such as GRPO.
- Experience with fine-tuning frameworks such as TRL, verl, OpenRLHF, or a custom training loop.
- Experience with fast rollouts using vLLM or SGLang.
- Experience with multi-GPU training using FSDP.
- Experience training tool-using or multi-turn agents.
- Experience with execution sandboxes, verifiers, evaluation harnesses, or tooling used by other engineers.
- Familiarity with open-weight model families such as Qwen or Llama.
- Familiarity with LoRA and similar techniques.

## Compensation and Benefits
- Competitive salary and equity package.
- Health, dental, and vision benefits.
- Flexible vacation policy.

## Similar jobs

- [Machine Learning Engineer III](https://hotfix.jobs/jobs/2359be26-5005-4fa7-94c9-8a86066a6bb5) - PathAI - Boston, MA - $131k – $200k/yr
- [Software Engineer, AI Platform](https://hotfix.jobs/jobs/7dcee5ac-38bb-4da0-b399-0b5896976722) - Fab2 - Austin, TX - $140k – $200k/yr
- [Member of Technical Staff - Image / Video Generation](https://hotfix.jobs/jobs/405a95a1-d8ff-49a2-9851-5a5e1cf05ad8) - Black Forest Labs - Freiburg, Germany - €130k – €340k/yr
- [Machine Learning Engineer](https://hotfix.jobs/jobs/41c0437e-b302-48c7-b2a7-5f426ba95cd8) - Mariana Minerals - Houston, TX - $140k – $180k/yr
- [Applied AI Engineer](https://hotfix.jobs/jobs/2f2e7873-cc6a-476d-a418-7b067cadf604) - Mintlify - San Francisco, CA - $130k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/8fc0d65b-35e7-4ce2-af19-eea2030c8d3d
**Canonical:** https://hotfix.jobs/jobs/8fc0d65b-35e7-4ce2-af19-eea2030c8d3d