# Senior Research Engineer, LLM Training & Post-Training

**Company:** [Lightning AI](https://hotfix.jobs/companies/lightning-ai)
**Location:** New York, NY, San Francisco, CA, Seattle, WA
**Role:** ML Engineering
**Salary:** $165k – $310k/yr
**Experience:** 5+ years
**Skills:** Python, PyTorch, LLMs, Transformers, Distributed Training, Multi-Gpu Systems, Supervised Fine-Tuning, RLHF, Dpo, Deepspeed, Fsdp, CUDA, Triton, Hugging Face Transformers, vLLM
**Posted:** 2026-08-12

> Build and optimize large language model training and post-training pipelines, improving model quality, distributed performance, evaluation, and production readiness. The role requires deep PyTorch and transformer experience, strong distributed-systems and software-engineering skills, and expertise in modern LLM optimization techniques.

## Job Description

## Responsibilities
- Design, build, and optimize training and post-training pipelines for large language models.
- Improve model quality through supervised fine-tuning, continued pretraining, preference optimization, reinforcement learning, evaluation, and experimentation.
- Build and improve PyTorch-based training infrastructure, tooling, and developer workflows.
- Optimize distributed training across multi-GPU environments by improving throughput, memory efficiency, scalability, and GPU utilization.
- Investigate challenging model training issues, including convergence, instability, communication overhead, and performance bottlenecks.
- Design evaluation methodologies, benchmark models, analyze failure modes, and guide model improvements through experimentation.
- Collaborate directly with customers to understand real-world workloads and translate those learnings into improvements across the research platform.
- Partner with research, infrastructure, and platform engineering teams to build production-ready AI systems.
- Contribute to open-source projects through new features, tooling improvements, documentation, and community engagement.

## Requirements
- Significant experience training, fine-tuning, evaluating, and optimizing transformer-based language models using PyTorch.
- Experience with modern LLM training and post-training techniques such as continued pretraining, SFT, RLHF, preference optimization, DPO, PPO, GRPO, and reward modeling.
- Strong understanding of distributed training and multi-GPU systems, with experience improving training performance, scalability, or efficiency.
- Strong software engineering fundamentals, including building production-quality Python software and research tooling.
- Experience designing experiments, evaluating model performance, and debugging complex training or optimization issues.
- Excellent communication and collaboration skills across research, product, infrastructure, and customer-facing engagements.
- Comfort working in fast-moving, ambiguous environments where priorities evolve over time.
- Master's degree, PhD, or equivalent industry experience in Machine Learning, AI, Computer Science, or a related field.

## Nice-to-Haves
- Experience with DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, or similar training frameworks.
- Experience with CUDA, Triton, vLLM, SGLang, TensorRT, or other AI systems and performance optimization technologies.
- GPU performance optimization, mixed precision, memory optimization, or distributed training optimization experience.
- Open-source contributions, research publications, or production AI platforms supporting large-scale training or inference workloads.
- Startup experience or experience working on highly cross-functional engineering teams.

## Compensation & Benefits
- Anticipated annual base salary: **$165,000–$310,000 USD**.
- Discretionary bonus and meaningful equity component.
- Medical, dental, and vision coverage for employees and eligible dependents.
- 401(k) matching and pension contributions where applicable.
- Unlimited PTO, company holidays, and floating holidays.
- Two-week company-wide winter break.
- Paid parental and family leave.
- Annual learning and development allowance.
- Wellness and work-from-home stipends.
- Four weeks of paid sabbatical after four years of service.
- Flexible schedules and a hybrid work model for office-based teams.
- Complimentary meals at office hubs.
- Benefits may vary by location, team, and role.

## Similar jobs

- [Senior Software Engineer - Machine Learning Platform](https://hotfix.jobs/jobs/da7a28b9-77d3-4a99-bd40-c739cdd55e8a) - Upstart - Remote - $167k – $230k/yr
- [Senior Robotics Software Engineer, Perception](https://hotfix.jobs/jobs/66889ba9-c493-4f25-ad8f-f79806865d00) - Parallel Systems - Los Angeles, CA - $163k – $212k/yr
- [Lead Applied AI Engineer](https://hotfix.jobs/jobs/e352036f-4480-44cd-a998-d33829aad30e) - LangChain - New York, NY - $160k – $200k/yr
- [Senior Data Science Engineer](https://hotfix.jobs/jobs/478eed0c-ff82-4528-822f-35cd9032cf1f) - LegitScript - Remote - $160k – $175k/yr
- [Senior Algorithm Engineer](https://hotfix.jobs/jobs/95f7bb2c-2fbf-4a7e-9c84-6633eef28354) - Beacon Biosignals - Remote - $170k – $190k/yr

**Apply:** https://hotfix.jobs/jobs/fd8081ee-45e7-4d2b-8a04-944f676278ac
**Canonical:** https://hotfix.jobs/jobs/fd8081ee-45e7-4d2b-8a04-944f676278ac