Skip to content
Together AITogether AISan Francisco, CA

Research Engineer, Large-Scale Training

Research Engineer turning efficient foundation model training research into robust high-performance production systems at Together AI. Optimize large-scale training infrastructure, profile bottlenecks, integrate new models, and productionize novel methods in close partnership with scientists.

200k – 290k/yr
On-siteML Engineering

About the role

Responsibilities

  • Design, implement, and optimize core components of Together's large-scale training infrastructure.
  • Integrate new model architectures, validate training correctness and convergence, and optimize performance for production fine-tuning workloads.
  • Profile distributed training workloads to identify and eliminate bottlenecks across compute, memory, and communication.
  • Design and execute experiments to validate performance hypotheses and benchmark new approaches against state-of-the-art methods.
  • Partner closely with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
  • Rapidly enable support for newly released open-source foundation models on the Together platform.
  • Build and maintain experimental infrastructure that accelerates research while ensuring production-quality reliability and scalability.

Requirements

  • Demonstrated ability to independently take ambiguous performance or infrastructure problems from investigation through deployment.
  • Strong programming skills in Python and PyTorch, with an emphasis on writing efficient, maintainable code.
  • Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node environments.
  • Solid understanding of ML systems fundamentals, including GPU architecture, mixed-precision training, and distributed training paradigms such as data, tensor, pipeline, or expert parallelism.
  • Strong communication skills and the ability to collaborate effectively with both researchers and engineers.
  • Passion for staying current with advances in AI research and applying them to real-world systems.
  • Excitement about translating cutting-edge research into production systems that deliver customer impact.

Nice to Have

  • Experience writing optimized NVIDIA GPU kernels using CUDA or Triton, or implementing communication collectives with technologies such as NCCL or NVSHMEM.
  • Experience with large-scale training frameworks such as FSDP, DeepSpeed, Megatron-LM, or custom distributed training systems.
  • Experience optimizing distributed training for compute efficiency, memory efficiency, or scalability.
  • Experience running and managing large-scale GPU experiments, including scheduling, monitoring, and fault tolerance.
  • Contributions to widely used open-source ML or ML systems projects.
  • Experience building or operating ML products or managed services used by external customers.

Compensation

US base salary range for this full-time position is $200,000 - $290,000.

Skills

PythonPyTorchCUDAtritonncclnvshmemfsdpdeepspeedmegatron-lmgpu architecturemixed-precision trainingDistributed Training

Similar roles

ML Engineering jobs
Cantina

Machine Learning Engineer - Voice Conversion

CantinaUnited States

Build and productionize large-scale generative speech models for voice conversion and related capabilities. The role combines research, data, evaluation, distributed training, performance optimization, and responsible deployment of speech systems.

200k – 220k/yrRemoteML Engineering
Siftstack

Software Engineer, Backend

SiftstackSan Francisco, CA +1

Build and operate dependable agentic AI systems that analyze large-scale hardware telemetry for aerospace and defense teams. Design tools, execution environments, distributed job systems on Kubernetes, and evaluation frameworks while owning the product end-to-end and speaking directly with customers.

200k – 250k/yrHybrid3+ YOEML Engineering
Console

Research Engineer

ConsoleSan Francisco, CA

Research Engineer building self-improving AI agent systems at Console. Develop eval/optimization loops, fine-tune specialist models, and improve agent reasoning over enterprise context using production data to drive measurable gains in quality, latency, and reliability.

200k – 350k/yrOn-siteML Engineering
Kepler

Machine Learning Engineer

KeplerNew York, NY

Build and own ML models, fine-tuning, evaluation harnesses, and routing for Kepler's AI agent harness in finance. Requires 5+ years production software experience and shipped ML systems focused on correctness, evals, and real-world reliability.

200k – 280k/yrOn-site5+ YOEML Engineering
AfterQuery

Machine Learning Engineer

AfterQuerySan Francisco, CA

Build production ML systems for measuring, predicting, and scaling data quality for frontier AI models. Requires 3-6 years experience in applied ML or related production systems (ranking, recommendations, data quality, fraud) plus strong software engineering skills.

200k – 300k/yrOn-site3+ YOEML Engineering