Research Engineer - Distributed Training
Research Engineer building and optimizing distributed infrastructure for frontier-scale model training and reinforcement learning. The role requires strong AI systems experience, PyTorch and distributed-training expertise, GPU performance optimization, and familiarity with parallelism and large-scale clusters.
About the job
Responsibilities
- Build and optimize distributed training infrastructure for pre-training and large-scale reinforcement learning workloads.
- Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.
- Design and implement low-level performance optimizations, including kernels, communication paths, and runtime improvements.
- Develop distributed training systems using data, tensor, and pipeline parallelism.
- Help shape reinforcement learning training architecture, including asynchronous rollout and post-training systems.
- Contribute to open-source libraries and internal infrastructure for frontier-scale model training.
- Collaborate with researchers and infrastructure engineers to translate bottlenecks into systems improvements.
- Work with training systems, inference systems, compiler/runtime tooling, and hardware-aware optimization techniques.
Requirements
- Strong systems engineering experience in AI/ML infrastructure, particularly large-scale model training or inference.
- Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, or Ray.
- Experience optimizing performance across kernels, memory movement, communication overhead, or parallelization strategies.
- Hands-on experience with data, tensor, and pipeline parallelism.
- Strong understanding of GPU architecture, profiling, and performance debugging.
- Ability to identify bottlenecks across the stack and drive improvements from first principles.
- Comfort working on ambiguous problems with high ownership.
Nice-to-haves
- Experience writing or optimizing CUDA or Triton kernels.
- Compiler or runtime optimization experience for ML systems.
- Experience with reinforcement learning infrastructure, rollout systems, or asynchronous training pipelines.
- Experience with multi-node GPU clusters and high-performance networking.
- Contributions to open-source ML systems or infrastructure projects.
- Interest in publishing technical work or writing engineering blogs.
Compensation
- Cash compensation range: $150,000–$350,000, plus equity incentives.
- Flexible work arrangements with options to work remotely or in person at offices in San Francisco.
- Visa sponsorship and relocation assistance for international candidates.
Skills
PyTorch, Pytorch Distributed, Deepspeed, Fsdp, Megatron, vLLM, Ray, CUDA, Triton, Gpu Architecture, Data Parallelism, Tensor Parallelism, Pipeline Parallelism, High-Performance Networking, Reinforcement Learning
Similar jobs
ML Engineering jobsBuild and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.
Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.
Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.
Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.