Technical Lead Manager - Training Runtime, Data Movement
Hands-on Technical Lead Manager to own dataset reads and data movement infrastructure for large-scale model training at OpenAI. Sets direction for APIs, storage contracts, versioning, reliability, and debugging tools across training frameworks.
About the job
Responsibilities
- Design and build a unified dataset read platform for multiple current and future training frameworks.
- Define dataset APIs, storage-format expectations, registration/versioning, and migration paths that make data access reproducible and maintainable.
- Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts.
- Build terminal and web-based visualizers that let teams inspect text, multimodal, and reinforcement learning data late in the pipeline.
- Write and review production code in core data loading, service, caching, and reliability paths.
- Partner with teams working on training frameworks, reinforcement learning, multimodal models, storage, runtime, and cluster infrastructure.
Requirements
- Have built or owned dataset, data loading, storage, or distributed training infrastructure at large scale (e.g. torch.utils.data).
- Care equally about API design, debugging ergonomics, performance, and bit-level correctness.
- Understand the failure modes of large distributed training jobs and know how data systems can create or prevent them.
- Have experience with stateful iterators, checkpoint/restart semantics, caching, remote services, or high-throughput storage reads.
- Are comfortable working across Python and lower-level systems code; Rust or C++ experience is useful but not required.
- Have worked with multimodal, video, reinforcement learning, or pretraining data pipelines where small data bugs are expensive and hard to diagnose.
- Can lead through code and technical judgment before a team exists, and can later manage engineers without losing the hands-on edge.
- Obsess over developer experience by eliminating friction, such as manual preprocessing scripts and niche cluster-specific bugs.
Skills
Python, Rust, C++, Torch.Utils.Data, Distributed Training, Data Loading, Stateful Iterators, Checkpoint/Restart, Caching, High-Throughput Storage
Similar jobs
Engineering Management jobsLeads a player-coach engineering team building reliable, scalable multimodal data products across image, video, audio, and document workflows. The role combines hands-on backend and distributed-systems architecture with team management, recruiting, and product delivery.
Leads the team and platform responsible for Gusto’s agentic coding environments, AI code review, and supporting CI and security infrastructure. The role requires substantial engineering leadership experience, hands-on use of agentic coding tools, vendor-contract ownership, and strong judgment on AI autonomy, risk, and infrastructure spending.
Leads an engineering team building AI-native internal productivity tools, partnering with business stakeholders and owning reliable cloud services. Requires 6+ years of engineering experience, 2+ years managing engineering teams, and experience with web applications, APIs, and cloud infrastructure.
Leads engineering teams building Garner’s core healthcare product, including data ingestion systems and new product experiences with potential AI enhancements. Requires 6+ years of engineering experience, 2+ years of engineering management, and familiarity with web applications, APIs, and AWS.
Leads and grows the Query Serving engineering team responsible for Mixpanel’s analytics entrypoint, query APIs, planning infrastructure, and scale protections. The role requires 6+ years of engineering experience, leadership experience, infrastructure expertise, and production experience with AI-powered products.