Research Engineer turning efficient foundation model training research into robust high-performance production systems at Together AI. Optimize large-scale training infrastructure, profile bottlenecks, integrate new models, and productionize novel methods in close partnership with scientists.
200k – 290k/yr
On-siteML Engineering
About the role
Responsibilities
Design, implement, and optimize core components of Together's large-scale training infrastructure.
Integrate new model architectures, validate training correctness and convergence, and optimize performance for production fine-tuning workloads.
Profile distributed training workloads to identify and eliminate bottlenecks across compute, memory, and communication.
Design and execute experiments to validate performance hypotheses and benchmark new approaches against state-of-the-art methods.
Partner closely with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
Rapidly enable support for newly released open-source foundation models on the Together platform.
Build and maintain experimental infrastructure that accelerates research while ensuring production-quality reliability and scalability.
Requirements
Demonstrated ability to independently take ambiguous performance or infrastructure problems from investigation through deployment.
Strong programming skills in Python and PyTorch, with an emphasis on writing efficient, maintainable code.
Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node environments.
Solid understanding of ML systems fundamentals, including GPU architecture, mixed-precision training, and distributed training paradigms such as data, tensor, pipeline, or expert parallelism.
Strong communication skills and the ability to collaborate effectively with both researchers and engineers.
Passion for staying current with advances in AI research and applying them to real-world systems.
Excitement about translating cutting-edge research into production systems that deliver customer impact.
Nice to Have
Experience writing optimized NVIDIA GPU kernels using CUDA or Triton, or implementing communication collectives with technologies such as NCCL or NVSHMEM.
Experience with large-scale training frameworks such as FSDP, DeepSpeed, Megatron-LM, or custom distributed training systems.
Experience optimizing distributed training for compute efficiency, memory efficiency, or scalability.
Experience running and managing large-scale GPU experiments, including scheduling, monitoring, and fault tolerance.
Contributions to widely used open-source ML or ML systems projects.
Experience building or operating ML products or managed services used by external customers.
Compensation
US base salary range for this full-time position is $200,000 - $290,000.
Skills
PythonPyTorchCUDAtritonncclnvshmemfsdpdeepspeedmegatron-lmgpu architecturemixed-precision trainingDistributed Training
Build and productionize large-scale generative speech models for voice conversion and related capabilities. The role combines research, data, evaluation, distributed training, performance optimization, and responsible deployment of speech systems.
200k – 220k/yrRemoteML Engineering
Software Engineer, Backend
SiftstackSan Francisco, CA +1
Build and operate dependable agentic AI systems that analyze large-scale hardware telemetry for aerospace and defense teams. Design tools, execution environments, distributed job systems on Kubernetes, and evaluation frameworks while owning the product end-to-end and speaking directly with customers.
200k – 250k/yrHybrid3+ YOEML Engineering
Research Engineer
ConsoleSan Francisco, CA
Research Engineer building self-improving AI agent systems at Console. Develop eval/optimization loops, fine-tune specialist models, and improve agent reasoning over enterprise context using production data to drive measurable gains in quality, latency, and reliability.
200k – 350k/yrOn-siteML Engineering
Machine Learning Engineer
KeplerNew York, NY
Build and own ML models, fine-tuning, evaluation harnesses, and routing for Kepler's AI agent harness in finance. Requires 5+ years production software experience and shipped ML systems focused on correctness, evals, and real-world reliability.
200k – 280k/yrOn-site5+ YOEML Engineering
Machine Learning Engineer
AfterQuerySan Francisco, CA
Build production ML systems for measuring, predicting, and scaling data quality for frontier AI models. Requires 3-6 years experience in applied ML or related production systems (ranking, recommendations, data quality, fraud) plus strong software engineering skills.