Research Engineer - RL Infrastructure
Build and optimize infrastructure for frontier-scale reinforcement learning and distributed model training, including kernels, runtimes, parallelism, and asynchronous rollouts. The role requires strong AI/ML systems experience, PyTorch expertise, and GPU performance optimization skills.
About the job
Responsibilities
- Build and optimize systems infrastructure for large-scale reinforcement learning and distributed training workloads in the
prime-rlframework. - Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.
- Design and implement low-level performance optimizations, including kernels, communication paths, and runtime improvements.
- Develop distributed training systems spanning data, tensor, and pipeline parallel workloads.
- Help shape RL training architecture, including asynchronous rollout and post-training systems.
- Contribute to open-source libraries and internal infrastructure for frontier-scale model training.
- Collaborate with researchers and infrastructure engineers to turn bottlenecks into systems improvements.
Requirements
- Strong systems engineering experience in AI/ML infrastructure, particularly large-scale model training or inference.
- Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, or Ray.
- Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategies.
- Hands-on experience with data parallelism, tensor parallelism, and pipeline parallelism.
- Strong understanding of GPU architecture, profiling, and performance debugging.
- Ability to identify bottlenecks across the stack and drive improvements from first principles.
- Comfort working in a fast-moving environment with ambiguous problems and high ownership.
Nice to Have
- Experience writing or optimizing CUDA or Triton kernels.
- Compiler or runtime optimization experience for ML systems.
- Experience with RL training infrastructure, rollout systems, or asynchronous training pipelines.
- Experience with multi-node GPU clusters and high-performance networking.
- Contributions to open-source ML systems or infrastructure projects.
- Interest in publishing technical work or sharing insights through engineering blogs and technical writing.
Compensation and Benefits
- Cash compensation of $150,000–$350,000 plus equity.
- Flexible work arrangements, with the option to work remotely or in person from the San Francisco office.
- Visa sponsorship and relocation support for international candidates.
- Quarterly team offsites, hackathons, conferences, and learning opportunities.
Skills
PyTorch, Pytorch Distributed, Deepspeed, Fsdp, Megatron, vLLM, Ray, CUDA, Triton, Gpu Architecture, Distributed Training, Reinforcement Learning, High-Performance Networking, Compiler Optimization, Linux
Similar jobs
ML Engineering jobsBuild and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.
Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.
Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.
Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.