Skip to content
OpenAIOpenAI

Software Engineer, Data Infrastructure

Builds and operates scalable data infrastructure including compute fleets, storage systems, and streaming platforms to support OpenAI's AI products, research, and analytics. Requires 4+ years in data or infrastructure engineering with expertise in Spark, Kafka, and distributed systems.

About the job

Responsibilities

  • Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security.
  • Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient.
  • Accelerate company productivity by empowering your fellow engineers & teammates with excellent data tooling and systems.
  • Collaborate with product, research and analytics teams to build the technical foundations capabilities that unlock new features and experiences.
  • Own the reliability of the systems you build, including participation in an on-call rotation for critical incidents.

Requirements

  • 4+ years in data infrastructure engineering OR 4+ years in infrastructure engineering with a strong interest in data.
  • Experience supporting Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms.
  • Well-versed in infrastructure tooling like Terraform.
  • Experienced in debugging large-scale distributed systems.

Nice-to-Haves

  • Comfortable with ambiguity and rapid change.
  • Intrinsic desire to learn and fill in missing skills, and talent for sharing learnings clearly.

Skills

Spark, Kafka, Flink, Airflow, Trino, Iceberg, Terraform, Delta Lake, Kubernetes, Chronon

Abridge

Abridge

San Francisco, CA

Data Engineer
$185k+/yrHybrid5+ YOEData Engineering

Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.

xAI

xAI

Palo Alto, CA

Analytics Engineer - X
$180k+/yrOn-site4+ YOEData Engineering

Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.

Vanta

Vanta

Remote

Operations Manager, Signal Systems
$176k+/yrRemoteData Engineering

Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.

Imprint

Imprint

New York, NY
Infrastructure Engineer
$170k+/yrOn-site5+ YOEData Engineering

Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Data Engineer - Axion
$200k+/yrOn-site5+ YOEData Engineering

Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.