Research Engineer, Post-Training Inference
Research Engineer building platforms to customize open-source LLMs via fine-tuning, RL, and evaluation. Focus on integrating post-training with inference engines (vLLM, SGLang, TensorRT-LLM), optimizing for RL workloads, and ensuring production reliability. Requires 2+ years ML production experience and strong Python/Go skills.
About the job
Responsibilities
- Design and build Together’s systems for customizing open-source models
- Build integrations between the Model Shaping and Inference platforms to ensure a seamless path from post-training to serving production workloads
- Add features to inference engines for large-scale post-training experiments, including optimizations for RL workloads
- Make sure the service is stable and robust, participating in an on-call rotation and ensuring 24/7 availability of our platform
Requirements
- 2+ years of experience building and deploying machine learning-based services in a production environment
- Hands-on experience with modern inference engines, such as SGLang, vLLM, and TensorRT-LLM
- Familiar with the latest methods for fine-tuning LLMs and other AI models
- Strong software engineering background in Python or Go
- Stay up to date with the latest advances and trends in the machine learning community
Nice-to-Haves
- Serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes
- Optimizing the performance of RL training workloads
- Developing CUDA/Triton/CuTE DSL kernels for inference
- Developing large-scale and high-load production systems
- Maintaining or contributing to open-source ML projects
- Managing machine learning workloads on Kubernetes clusters
Compensation
US base salary range for this full-time position is $200,000 - $290,000.
Skills
Sglang, vLLM, Tensorrt-Llm, Python, Go, CUDA, Triton, Kubernetes, Lora, RLHF, Llm Fine-Tuning
Similar jobs
ML Engineering jobsResearch Engineer focused on building and deploying real-time audio and speech models for conversational voice agents. The role requires experience with speech or multimodal machine learning, production inference, Python, and PyTorch, with emphasis on taking research from prototype to measurable production impact.
Research Engineer focused on making conversational AI agents safe, reliable, and controllable in production. The role develops evaluations, safeguards, post-training methods, and monitoring systems, requiring 2+ years of AI/ML or safety experience and strong Python and production engineering skills.
Research Engineer designing post-training infrastructure and running controlled experiments to measure how datasets affect foundation-model behavior. Requires at least 2 years of ML or research engineering experience, strong Python, and hands-on experience with PyTorch, JAX, Ray, Slurm, and LLM post-training.
Build and productionize scalable machine learning models and systems for underwriting and portfolio management. The role requires a bachelor's degree and at least two years of experience shipping ML systems, plus expertise in model development, deployment, data pipelines, and deep learning.
Machine learning engineer who trains, evaluates, and productionizes models and LLM-powered applications for financial products. Requires 2+ years of ML systems experience, strong Python and PyTorch skills, production data pipelines, model evaluation, and API development.