Skip to content

Senior Research Engineer, LLM Training & Post-Training

Build and optimize large language model training and post-training pipelines, improving model quality, distributed performance, evaluation, and production readiness. The role requires deep PyTorch and transformer experience, strong distributed-systems and software-engineering skills, and expertise in modern LLM optimization techniques.

About the job

Responsibilities

  • Design, build, and optimize training and post-training pipelines for large language models.
  • Improve model quality through supervised fine-tuning, continued pretraining, preference optimization, reinforcement learning, evaluation, and experimentation.
  • Build and improve PyTorch-based training infrastructure, tooling, and developer workflows.
  • Optimize distributed training across multi-GPU environments by improving throughput, memory efficiency, scalability, and GPU utilization.
  • Investigate challenging model training issues, including convergence, instability, communication overhead, and performance bottlenecks.
  • Design evaluation methodologies, benchmark models, analyze failure modes, and guide model improvements through experimentation.
  • Collaborate directly with customers to understand real-world workloads and translate those learnings into improvements across the research platform.
  • Partner with research, infrastructure, and platform engineering teams to build production-ready AI systems.
  • Contribute to open-source projects through new features, tooling improvements, documentation, and community engagement.

Requirements

  • Significant experience training, fine-tuning, evaluating, and optimizing transformer-based language models using PyTorch.
  • Experience with modern LLM training and post-training techniques such as continued pretraining, SFT, RLHF, preference optimization, DPO, PPO, GRPO, and reward modeling.
  • Strong understanding of distributed training and multi-GPU systems, with experience improving training performance, scalability, or efficiency.
  • Strong software engineering fundamentals, including building production-quality Python software and research tooling.
  • Experience designing experiments, evaluating model performance, and debugging complex training or optimization issues.
  • Excellent communication and collaboration skills across research, product, infrastructure, and customer-facing engagements.
  • Comfort working in fast-moving, ambiguous environments where priorities evolve over time.
  • Master's degree, PhD, or equivalent industry experience in Machine Learning, AI, Computer Science, or a related field.

Nice-to-Haves

  • Experience with DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, or similar training frameworks.
  • Experience with CUDA, Triton, vLLM, SGLang, TensorRT, or other AI systems and performance optimization technologies.
  • GPU performance optimization, mixed precision, memory optimization, or distributed training optimization experience.
  • Open-source contributions, research publications, or production AI platforms supporting large-scale training or inference workloads.
  • Startup experience or experience working on highly cross-functional engineering teams.

Compensation & Benefits

  • Anticipated annual base salary: $165,000–$310,000 USD.
  • Discretionary bonus and meaningful equity component.
  • Medical, dental, and vision coverage for employees and eligible dependents.
  • 401(k) matching and pension contributions where applicable.
  • Unlimited PTO, company holidays, and floating holidays.
  • Two-week company-wide winter break.
  • Paid parental and family leave.
  • Annual learning and development allowance.
  • Wellness and work-from-home stipends.
  • Four weeks of paid sabbatical after four years of service.
  • Flexible schedules and a hybrid work model for office-based teams.
  • Complimentary meals at office hubs.
  • Benefits may vary by location, team, and role.

Skills

Python, PyTorch, LLMs, Transformers, Distributed Training, Multi-Gpu Systems, Supervised Fine-Tuning, RLHF, Dpo, Deepspeed, Fsdp, CUDA, Triton, Hugging Face Transformers, vLLM

Upstart

Upstart

United States

Senior Software Engineer - Machine Learning Platform
$167k+/yrRemote6+ YOEML Engineering

Build and operate backend infrastructure for machine learning model training, serving, feature management, and marketplace simulation. The role requires 6+ years of software engineering experience, distributed systems expertise, and experience with production ML platforms.

Parallel Systems

Parallel Systems

Los Angeles, CA

Senior Robotics Software Engineer, Perception
$163k+/yrOn-site5+ YOEML Engineering

Develop and deploy real-time perception and sensor-fusion software for autonomous battery-electric rail vehicles. The role requires strong robotics, geometry-based computer vision, C/C++ and Rust experience, plus hands-on work with multimodal sensors and production systems.

LangChain

LangChain

New York, NY

Lead Applied AI Engineer
$160k+/yrOn-site7+ YOEML Engineering

Leads hands-on development and deployment of production AI agents and workflows, sets technical standards, and mentors engineers. Requires 4+ years of software engineering experience, production LLM application experience, and strong Python or TypeScript skills.

LegitScript

LegitScript

United States

Senior Data Science Engineer
$160k+/yrRemote5+ YOEML Engineering

Owns the full lifecycle of data and ML solutions, from ingestion and feature-ready datasets through production deployment and business-impact measurement. The role combines data engineering, applied machine learning, MLOps, and generative AI to build risk detection capabilities.

Beacon Biosignals

Beacon Biosignals

United States

Senior Algorithm Engineer
$170k+/yrRemote5+ YOEML Engineering

The Senior Algorithm Engineer leads development and production deployment of machine and deep learning algorithms for biosignal and medical-device applications. The role requires 5+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated health or similar domains.