Software Engineer, Data Infrastructure - Research
Designs and implements dataset infrastructure for OpenAI's large-scale LLM training stack, including standardized APIs for multimodal data, scaling pipelines across GPU fleets, and performance debugging. Requires strong distributed systems experience and collaboration with researchers.
About the job
Responsibilities
- Design and maintain standardized dataset APIs, including for multimodal (MM) data that cannot fit in memory.
- Build proactive testing and scale validation pipelines for dataset loading at GPU scale.
- Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience.
- Document and maintain dataset interfaces so they are discoverable, consistent, and easy for other teams to adopt.
- Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged once standardized.
- Debug and resolve performance bottlenecks in distributed dataset loading (e.g., straggler systems slowing global training).
- Provide visualization and inspection tools to surface errors, bugs, or bottlenecks in datasets.
Requirements
- Strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.
- Experience building APIs, modular code, and scalable abstractions, while recognizing that abstractions ultimately serve the users and UX is an important part of the abstractions design.
- Comfortable debugging bottlenecks across large fleets of machines.
- Take pride in building infrastructure that “just works,” and find joy in being the guardian of reliability and scale.
- Collaborative, humble, and excited to own a foundational (if not glamorous) part of the ML stack.
Nice-to-Haves
- Background knowledge in data math, probability, or distributed data theory.
- Worked with GPU-scale distributed systems or dataset scaling for real-time data.
Skills
Distributed Systems, Data Pipelines, APIs, GPU, PyTorch, Kubernetes, Python, Rust, Scalable Abstractions, Dataset Loading
Similar jobs
Data Engineering jobsOwn end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.
Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.
Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.
Build and evolve reliable analytics infrastructure, pipelines, schemas, and foundational datasets supporting quantitative research across strategies. The role requires strong Python and SQL skills, distributed data-platform experience, and ownership of observability, performance, and reproducibility.
Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.