Skip to content
ExaExa

Software Engineer, Distributed Data Systems

Architects and builds massive-scale data infrastructure for web crawling, embedding model training, and real-time search, handling hundreds of petabytes. Requires expertise in lakehouse architectures, distributed processing pipelines, and streaming systems like Kafka and Flink.

About the job

Responsibilities

  • Architect and build data infrastructure for crawling billions of pages, training embedding models, and serving real-time search.
  • Design systems that scale to hundreds of petabytes.
  • Example projects: Design lakehouse architecture for 100+ PB web crawl data; build streaming pipelines processing billions of documents per day; architect data layer for embedding training on Ray; scale ClickHouse for petabyte-scale analytical queries.

Requirements

  • Deep understanding of lakehouse architectures (Delta Lake, Iceberg, Hudi) and when to use them.
  • Experience building and operating large-scale distributed data processing pipelines.
  • Hands-on experience with streaming data systems (Kafka, Flink, or similar).
  • Familiarity with Ray, Spark, or ClickHouse at production scale.
  • Obsessive focus on reliability and building systems that don't page you at 3am.

Nice-to-Haves

  • Experience with Lance or other vector-native storage formats.
  • Background in GPU-accelerated data processing (RAPIDS, cuDF).

Skills

Delta Lake, Iceberg, Hudi, Kafka, Flink, Ray, Spark, ClickHouse, Lance, Rapids, Cudf

xAI

xAI

Palo Alto, CA

Analytics Engineer - X
$180k+/yrOn-site4+ YOEData Engineering

Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.

Vanta

Vanta

Remote

Operations Manager, Signal Systems
$176k+/yrRemoteData Engineering

Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.

Abridge

Abridge

San Francisco, CA

Data Engineer
$185k+/yrHybrid5+ YOEData Engineering

Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.

Imprint

Imprint

New York, NY
Infrastructure Engineer
$170k+/yrOn-site5+ YOEData Engineering

Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Data Engineer - Axion
$200k+/yrOn-site5+ YOEData Engineering

Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.