Skip to content

Research Engineer, Post-Training Inference

Research Engineer building platforms to customize open-source LLMs via fine-tuning, RL, and evaluation. Focus on integrating post-training with inference engines (vLLM, SGLang, TensorRT-LLM), optimizing for RL workloads, and ensuring production reliability. Requires 2+ years ML production experience and strong Python/Go skills.

About the job

Responsibilities

  • Design and build Together’s systems for customizing open-source models
  • Build integrations between the Model Shaping and Inference platforms to ensure a seamless path from post-training to serving production workloads
  • Add features to inference engines for large-scale post-training experiments, including optimizations for RL workloads
  • Make sure the service is stable and robust, participating in an on-call rotation and ensuring 24/7 availability of our platform

Requirements

  • 2+ years of experience building and deploying machine learning-based services in a production environment
  • Hands-on experience with modern inference engines, such as SGLang, vLLM, and TensorRT-LLM
  • Familiar with the latest methods for fine-tuning LLMs and other AI models
  • Strong software engineering background in Python or Go
  • Stay up to date with the latest advances and trends in the machine learning community

Nice-to-Haves

  • Serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes
  • Optimizing the performance of RL training workloads
  • Developing CUDA/Triton/CuTE DSL kernels for inference
  • Developing large-scale and high-load production systems
  • Maintaining or contributing to open-source ML projects
  • Managing machine learning workloads on Kubernetes clusters

Compensation

US base salary range for this full-time position is $200,000 - $290,000.

Skills

Sglang, vLLM, Tensorrt-Llm, Python, Go, CUDA, Triton, Kubernetes, Lora, RLHF, Llm Fine-Tuning

Decagon

Decagon

San Francisco, CA
Research Engineer, Audio and Speech
$200k+/yrOn-site2+ YOEML Engineering

Research Engineer focused on building and deploying real-time audio and speech models for conversational voice agents. The role requires experience with speech or multimodal machine learning, production inference, Python, and PyTorch, with emphasis on taking research from prototype to measurable production impact.

Decagon

Decagon

San Francisco, CA
Research Engineer, Safety
$200k+/yrOn-site2+ YOEML Engineering

Research Engineer focused on making conversational AI agents safe, reliable, and controllable in production. The role develops evaluations, safeguards, post-training methods, and monitoring systems, requiring 2+ years of AI/ML or safety experience and strong Python and production engineering skills.

AfterQuery

AfterQuery

San Francisco, CA

Research Engineer
$210k+/yrOn-site2+ YOEML Engineering

Research Engineer designing post-training infrastructure and running controlled experiments to measure how datasets affect foundation-model behavior. Requires at least 2 years of ML or research engineering experience, strong Python, and hands-on experience with PyTorch, JAX, Ray, Slurm, and LLM post-training.

Stripe

Stripe

South San Francisco, CA

Machine Learning Engineer
$212k+/yrHybrid2+ YOEML Engineering

Build and productionize scalable machine learning models and systems for underwriting and portfolio management. The role requires a bachelor's degree and at least two years of experience shipping ML systems, plus expertise in model development, deployment, data pipelines, and deep learning.

Earnin

Earnin

Mountain View, CA

Machine Learning Engineer
$187k+/yrHybrid2+ YOEML Engineering

Machine learning engineer who trains, evaluates, and productionizes models and LLM-powered applications for financial products. Requires 2+ years of ML systems experience, strong Python and PyTorch skills, production data pipelines, model evaluation, and API development.