Research Scientist / Engineer – Training Infrastructure
Builds and optimizes distributed training infrastructure for large-scale multimodal AI models across thousands of GPUs. Requires deep expertise in PyTorch, CUDA, parallelization techniques, and GPU clusters.
About the job
Responsibilities
- Design, implement, and optimize efficient distributed training systems for models with thousands of GPUs
- Research and implement advanced parallelization techniques (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel)
- Build monitoring, visualization, and debugging tools for large-scale training runs
- Optimize training stability, convergence, and resource utilization across massive clusters
Experience
- Extensive experience with distributed PyTorch training and parallelisms in foundation model training
- Deep understanding of GPU clusters, networking, and storage systems
- Familiarity with communication libraries (NCCL, MPI) and distributed system optimization
(Preferred)
- Strong Linux systems administration and scripting capabilities
- Experience managing training runs across >100 GPUs
- Experience with containerization, orchestration, and cloud infrastructure
Compensation
Base pay range: $187,500 – $395,000 per year
Skills
PyTorch, CUDA, Distributed Systems, Fsdp, Tensor Parallel, Pipeline Parallel, Expert Parallel, Nccl, Mpi, Kubernetes
Similar jobs
ML Engineering jobsBuild production-grade AI agents, evaluation infrastructure, and developer tooling that make AI-assisted engineering faster, safer, and reusable across teams. The role requires software engineering experience, platform or internal developer-product experience, and hands-on expertise with LLM integration and orchestration.
Build trustworthy infrastructure for production LLM agents, closed-loop evaluation, and autonomous research workflows. The role requires strong Python and distributed-systems experience, hands-on LLM post-training and inference knowledge, and experience operating agent systems at scale.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.
Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.