Research Engineer, Infrastructure
Builds and owns distributed training infrastructure, experiment orchestration, data pipelines, and performance optimizations for large-scale AI research on GPU clusters. Requires deep systems expertise, Python/C++/PyTorch proficiency, and ML understanding to accelerate frontier research.
About the job
Responsibilities
- Build and own distributed training infrastructure for large-scale GPU clusters, including job launchers, checkpointing, recovery, fault tolerance, and monitoring.
- Own infrastructure for scaling agent rollouts in VM sandboxes at RL training scales.
- Profile and optimize training throughput: data loading, communication, memory, compute efficiency to improve step time and MFU.
- Design experiment orchestration and tooling to launch, track, and analyze experiments.
- Build high-throughput, reliable data pipelines for training and evaluation.
- Debug and resolve training failures across GPUs, networking, numerics, and data.
- Implement and optimize parallelism strategies: data, tensor, pipeline, sequence.
- Anticipate research needs and build scaling infrastructure proactively.
Requirements
- Deep experience building/operating distributed training systems for large models.
- Strong systems engineering: distributed systems, networking, storage, performance reasoning.
- Proficiency in Python, C++; systems-level PyTorch or equivalent.
- Hands-on GPU performance profiling, memory optimization, compute efficiency.
- Experience with parallelism strategies for large model training.
- Track record building tooling to accelerate research workflows.
- Strong debugging in complex distributed systems.
- ML knowledge to engage with researchers.
Skills
Python, C++, PyTorch, Distributed Systems, Gpu Programming, Parallelism, Data Pipelines, Experiment Orchestration, Performance Profiling, Networking
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.