# Research Engineer - Distributed Training

**Company:** [Prime Intellect](https://hotfix.jobs/companies/prime-intellect)
**Location:** Remote
**Role:** ML Engineering
**Salary:** $150k – $350k/yr
**Skills:** PyTorch, Pytorch Distributed, Deepspeed, Fsdp, Megatron, vLLM, Ray, CUDA, Triton, Gpu Architecture, Data Parallelism, Tensor Parallelism, Pipeline Parallelism, High-Performance Networking, Reinforcement Learning
**Posted:** 2026-07-08

> Research Engineer building and optimizing distributed infrastructure for frontier-scale model training and reinforcement learning. The role requires strong AI systems experience, PyTorch and distributed-training expertise, GPU performance optimization, and familiarity with parallelism and large-scale clusters.

## Job Description

## Responsibilities
- Build and optimize distributed training infrastructure for pre-training and large-scale reinforcement learning workloads.
- Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.
- Design and implement low-level performance optimizations, including kernels, communication paths, and runtime improvements.
- Develop distributed training systems using data, tensor, and pipeline parallelism.
- Help shape reinforcement learning training architecture, including asynchronous rollout and post-training systems.
- Contribute to open-source libraries and internal infrastructure for frontier-scale model training.
- Collaborate with researchers and infrastructure engineers to translate bottlenecks into systems improvements.
- Work with training systems, inference systems, compiler/runtime tooling, and hardware-aware optimization techniques.

## Requirements
- Strong systems engineering experience in AI/ML infrastructure, particularly large-scale model training or inference.
- Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, or Ray.
- Experience optimizing performance across kernels, memory movement, communication overhead, or parallelization strategies.
- Hands-on experience with data, tensor, and pipeline parallelism.
- Strong understanding of GPU architecture, profiling, and performance debugging.
- Ability to identify bottlenecks across the stack and drive improvements from first principles.
- Comfort working on ambiguous problems with high ownership.

## Nice-to-haves
- Experience writing or optimizing CUDA or Triton kernels.
- Compiler or runtime optimization experience for ML systems.
- Experience with reinforcement learning infrastructure, rollout systems, or asynchronous training pipelines.
- Experience with multi-node GPU clusters and high-performance networking.
- Contributions to open-source ML systems or infrastructure projects.
- Interest in publishing technical work or writing engineering blogs.

## Compensation
- Cash compensation range: $150,000–$350,000, plus equity incentives.
- Flexible work arrangements with options to work remotely or in person at offices in San Francisco.
- Visa sponsorship and relocation assistance for international candidates.

## Similar jobs

- [AI Engineer, Enablement](https://hotfix.jobs/jobs/6ae315a1-a605-45dc-b1ee-88fa8e9dee24) - LangChain - New York, NY - $150k – $195k/yr
- [Member of Technical Staff — Frontier Data](https://hotfix.jobs/jobs/20efe17f-aa7c-401b-8add-7086d4571fba) - Roboflow - Remote - $150k – $300k/yr
- [Algorithm Engineer](https://hotfix.jobs/jobs/3bef67e8-d4e7-4866-a9b8-6c7863fe2e96) - Beacon Biosignals - Remote - $150k – $170k/yr
- [Software Engineer - Prediction and Planning ML](https://hotfix.jobs/jobs/21b9c778-e1ae-4695-b26d-fec68ea8a8cc) - Applied Intuition - Sunnyvale, CA - $151k – $258k/yr
- [Software Engineer, AI Platform](https://hotfix.jobs/jobs/7dcee5ac-38bb-4da0-b399-0b5896976722) - Fab2 - Austin, TX - $140k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/0429b6c1-b96b-4161-ac6a-ab224a9609de
**Canonical:** https://hotfix.jobs/jobs/0429b6c1-b96b-4161-ac6a-ab224a9609de