Skip to content
OpenAIOpenAI

Software Engineer, Distributed Data Systems (Sora)

Designs and scales distributed data infrastructure for large-scale multimodal training and evaluation at OpenAI. Collaborates with researchers to build reliable, high-performance systems handling massive data volumes in a fast-paced environment.

About the job

Responsibilities

  • Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security.
  • Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient.
  • Partner with researchers to deeply understand requirements and translate them into production-ready systems.
  • Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation.

Requirements

  • Strong experience with distributed systems and large-scale infrastructure with a strong interest in data.
  • Detail-oriented and bring rigor to building and maintaining reliable systems.
  • Excellent software engineering fundamentals and organizational skills.
  • Comfortable with ambiguity and rapid change.

Skills

Distributed Systems, Data Orchestration, Distributed Storage, Streaming Infrastructure, Machine Learning Infrastructure, Kubernetes, Spark, Apache Kafka, AWS, GCP

OpenAI

OpenAI

Mountain View, CA
Data Engineer, Monetization Data Platform
$230k+/yrOn-siteData Engineering

Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.

The Voleon Group

The Voleon Group

New York, NY
Software Engineer, Strategy Research Analytics
$230k+/yrRemote3+ YOEData Engineering

Build and evolve reliable analytics infrastructure, pipelines, schemas, and foundational datasets supporting quantitative research across strategies. The role requires strong Python and SQL skills, distributed data-platform experience, and ownership of observability, performance, and reproducibility.

Anyscale

Anyscale

San Francisco, CA

Software Engineer
$215k+/yrOn-site3+ YOEData Engineering

Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Data Operations
$250k+/yrHybridData Engineering

Own end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Data Engineer - Axion
$200k+/yrOn-site5+ YOEData Engineering

Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.