Senior Research Engineer, LLM Training & Post-Training
Build and optimize large language model training and post-training pipelines, improving model quality, distributed performance, evaluation, and production readiness. The role requires deep PyTorch and transformer experience, strong distributed-systems and software-engineering skills, and expertise in modern LLM optimization techniques.
About the job
Responsibilities
- Design, build, and optimize training and post-training pipelines for large language models.
- Improve model quality through supervised fine-tuning, continued pretraining, preference optimization, reinforcement learning, evaluation, and experimentation.
- Build and improve PyTorch-based training infrastructure, tooling, and developer workflows.
- Optimize distributed training across multi-GPU environments by improving throughput, memory efficiency, scalability, and GPU utilization.
- Investigate challenging model training issues, including convergence, instability, communication overhead, and performance bottlenecks.
- Design evaluation methodologies, benchmark models, analyze failure modes, and guide model improvements through experimentation.
- Collaborate directly with customers to understand real-world workloads and translate those learnings into improvements across the research platform.
- Partner with research, infrastructure, and platform engineering teams to build production-ready AI systems.
- Contribute to open-source projects through new features, tooling improvements, documentation, and community engagement.
Requirements
- Significant experience training, fine-tuning, evaluating, and optimizing transformer-based language models using PyTorch.
- Experience with modern LLM training and post-training techniques such as continued pretraining, SFT, RLHF, preference optimization, DPO, PPO, GRPO, and reward modeling.
- Strong understanding of distributed training and multi-GPU systems, with experience improving training performance, scalability, or efficiency.
- Strong software engineering fundamentals, including building production-quality Python software and research tooling.
- Experience designing experiments, evaluating model performance, and debugging complex training or optimization issues.
- Excellent communication and collaboration skills across research, product, infrastructure, and customer-facing engagements.
- Comfort working in fast-moving, ambiguous environments where priorities evolve over time.
- Master's degree, PhD, or equivalent industry experience in Machine Learning, AI, Computer Science, or a related field.
Nice-to-Haves
- Experience with DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, or similar training frameworks.
- Experience with CUDA, Triton, vLLM, SGLang, TensorRT, or other AI systems and performance optimization technologies.
- GPU performance optimization, mixed precision, memory optimization, or distributed training optimization experience.
- Open-source contributions, research publications, or production AI platforms supporting large-scale training or inference workloads.
- Startup experience or experience working on highly cross-functional engineering teams.
Compensation & Benefits
- Anticipated annual base salary: $165,000–$310,000 USD.
- Discretionary bonus and meaningful equity component.
- Medical, dental, and vision coverage for employees and eligible dependents.
- 401(k) matching and pension contributions where applicable.
- Unlimited PTO, company holidays, and floating holidays.
- Two-week company-wide winter break.
- Paid parental and family leave.
- Annual learning and development allowance.
- Wellness and work-from-home stipends.
- Four weeks of paid sabbatical after four years of service.
- Flexible schedules and a hybrid work model for office-based teams.
- Complimentary meals at office hubs.
- Benefits may vary by location, team, and role.
Skills
Python, PyTorch, LLMs, Transformers, Distributed Training, Multi-Gpu Systems, Supervised Fine-Tuning, RLHF, Dpo, Deepspeed, Fsdp, CUDA, Triton, Hugging Face Transformers, vLLM
Similar jobs
ML Engineering jobsBuild and operate backend infrastructure for machine learning model training, serving, feature management, and marketplace simulation. The role requires 6+ years of software engineering experience, distributed systems expertise, and experience with production ML platforms.
Develop and deploy real-time perception and sensor-fusion software for autonomous battery-electric rail vehicles. The role requires strong robotics, geometry-based computer vision, C/C++ and Rust experience, plus hands-on work with multimodal sensors and production systems.
Leads hands-on development and deployment of production AI agents and workflows, sets technical standards, and mentors engineers. Requires 4+ years of software engineering experience, production LLM application experience, and strong Python or TypeScript skills.
Owns the full lifecycle of data and ML solutions, from ingestion and feature-ready datasets through production deployment and business-impact measurement. The role combines data engineering, applied machine learning, MLOps, and generative AI to build risk detection capabilities.
The Senior Algorithm Engineer leads development and production deployment of machine and deep learning algorithms for biosignal and medical-device applications. The role requires 5+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated health or similar domains.