Skip to content

Latest ML Engineering jobs at Together AI

Search
Location
8 jobs

Job results

Together AI

Research Engineer, Large-Scale Training

Together AISan Francisco, CA

Research Engineer turning efficient foundation model training research into robust high-performance production systems at Together AI. Optimize large-scale training infrastructure, profile bottlenecks, integrate new models, and productionize novel methods in close partnership with scientists.

200k – 290k/yrOn-siteML Engineering
Together AI

Research Engineer, Post-Training Inference

Together AISan Francisco, CA

Research Engineer building platforms to customize open-source LLMs via fine-tuning, RL, and evaluation. Focus on integrating post-training with inference engines (vLLM, SGLang, TensorRT-LLM), optimizing for RL workloads, and ensuring production reliability. Requires 2+ years ML production experience and strong Python/Go skills.

200k – 290k/yrOn-site2+ YOEML Engineering
Together AI

Research Intern, Model Shaping

Together AISan Francisco, CA

Research intern on the Model Shaping team working on post-training methods, efficient neural network training, and foundation model evaluation. Requires strong ML fundamentals and PyTorch/JAX experience.

121k – 131k/yrOn-siteEntry levelML Engineering
Together AI

Systems Research Engineer Intern - GPU Programming

Together AISan Francisco, CA

Intern developing and optimizing GPU-accelerated kernels for ML/AI applications. Requires strong GPU programming background (CUDA/Triton) and knowledge of performance optimization.

121k – 131k/yrOn-siteEntry levelML Engineering
Together AI

Staff Machine Learning Engineer, Voice AI

Together AISan Francisco, CA

Staff ML Engineer to own the model serving stack for real-time voice inference (STT, TTS, speech-to-speech) on H100/H200 GPUs. Drive latency/throughput optimization using TRT-LLM and SGLang for models like Whisper and Parakeet.

220k – 280k/yrOn-site8+ YOEML Engineering
Together AI

Senior Machine Learning Engineer, Voice AI

Together AISan Francisco, CA

Senior ML Engineer optimizes inference for voice AI models (STT, TTS, speech-to-speech) using engines like TensorRT-LLM and SGLang on GPUs. Requires 5+ years ML engineering with serving/inference expertise, Python/PyTorch proficiency, and production ML experience.

200k – 260k/yrOn-site5+ YOEML Engineering
Together AI

Research Engineer, Core ML

Together AISan Francisco, CA

Research Engineer building production ML systems at the intersection of efficient inference, RL/post-training, and serving engines. Translates algorithms into scalable infrastructure improving latency, throughput, and model quality. Requires 3+ years ML systems experience and advanced degree.

200k – 280k/yrOn-site3+ YOEML Engineering
Together AI

Machine Learning Engineer - Inference

Together AISan Francisco, CA

Optimizes and builds production inference systems for large language models at scale using PyTorch and high-performance tooling. Requires 3+ years experience in production code, OS concepts, and AI inference systems.

160k – 230k/yrOn-site3+ YOEML Engineering