# Research Engineer, Large-Scale Training

**Company:** [Together AI](https://hotfix.jobs/companies/together-ai)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $200k – $290k/yr
**Skills:** Python, PyTorch, CUDA, triton, nccl, nvshmem, fsdp, deepspeed, megatron-lm, gpu architecture, mixed-precision training, Distributed Training
**Posted:** 2026-07-30

> Research Engineer turning efficient foundation model training research into robust high-performance production systems at Together AI. Optimize large-scale training infrastructure, profile bottlenecks, integrate new models, and productionize novel methods in close partnership with scientists.

## Job Description

## Responsibilities
- Design, implement, and optimize core components of Together's large-scale training infrastructure.
- Integrate new model architectures, validate training correctness and convergence, and optimize performance for production fine-tuning workloads.
- Profile distributed training workloads to identify and eliminate bottlenecks across compute, memory, and communication.
- Design and execute experiments to validate performance hypotheses and benchmark new approaches against state-of-the-art methods.
- Partner closely with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
- Rapidly enable support for newly released open-source foundation models on the Together platform.
- Build and maintain experimental infrastructure that accelerates research while ensuring production-quality reliability and scalability.

## Requirements
- Demonstrated ability to independently take ambiguous performance or infrastructure problems from investigation through deployment.
- Strong programming skills in Python and PyTorch, with an emphasis on writing efficient, maintainable code.
- Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node environments.
- Solid understanding of ML systems fundamentals, including GPU architecture, mixed-precision training, and distributed training paradigms such as data, tensor, pipeline, or expert parallelism.
- Strong communication skills and the ability to collaborate effectively with both researchers and engineers.
- Passion for staying current with advances in AI research and applying them to real-world systems.
- Excitement about translating cutting-edge research into production systems that deliver customer impact.

## Nice to Have
- Experience writing optimized NVIDIA GPU kernels using CUDA or Triton, or implementing communication collectives with technologies such as NCCL or NVSHMEM.
- Experience with large-scale training frameworks such as FSDP, DeepSpeed, Megatron-LM, or custom distributed training systems.
- Experience optimizing distributed training for compute efficiency, memory efficiency, or scalability.
- Experience running and managing large-scale GPU experiments, including scheduling, monitoring, and fault tolerance.
- Contributions to widely used open-source ML or ML systems projects.
- Experience building or operating ML products or managed services used by external customers.

## Compensation
US base salary range for this full-time position is $200,000 - $290,000.

## Similar roles

- [Machine Learning Engineer - Voice Conversion](https://hotfix.jobs/jobs/08127c48-42e9-4d53-843a-3f2829ac5085) - Cantina - Remote - $200k – $220k/yr
- [Software Engineer, Backend](https://hotfix.jobs/jobs/5f7dd806-8fd4-4617-99d5-e6e01ce8ae71) - Siftstack - San Francisco, CA - $200k – $250k/yr
- [Research Engineer](https://hotfix.jobs/jobs/b4d18018-e64d-4828-a8b4-303225328e17) - Console - San Francisco, CA - $200k – $350k/yr
- [Machine Learning Engineer](https://hotfix.jobs/jobs/1b64040d-77ad-441e-85bb-9de3583cf9b8) - Kepler - New York, NY - $200k – $280k/yr
- [Machine Learning Engineer](https://hotfix.jobs/jobs/bc31e7f5-d976-4a97-a1be-bee1a0b92f94) - AfterQuery - San Francisco, CA - $200k – $300k/yr

**Apply:** https://hotfix.jobs/jobs/2a0f5a2d-314c-41f9-8731-308f3b6a53ba
**Canonical:** https://hotfix.jobs/jobs/2a0f5a2d-314c-41f9-8731-308f3b6a53ba