Software Engineer, Distributed Data Systems
Architects and builds massive-scale data infrastructure for web crawling, embedding model training, and real-time search, handling hundreds of petabytes. Requires expertise in lakehouse architectures, distributed processing pipelines, and streaming systems like Kafka and Flink.
About the job
Responsibilities
- Architect and build data infrastructure for crawling billions of pages, training embedding models, and serving real-time search.
- Design systems that scale to hundreds of petabytes.
- Example projects: Design lakehouse architecture for 100+ PB web crawl data; build streaming pipelines processing billions of documents per day; architect data layer for embedding training on Ray; scale ClickHouse for petabyte-scale analytical queries.
Requirements
- Deep understanding of lakehouse architectures (Delta Lake, Iceberg, Hudi) and when to use them.
- Experience building and operating large-scale distributed data processing pipelines.
- Hands-on experience with streaming data systems (Kafka, Flink, or similar).
- Familiarity with Ray, Spark, or ClickHouse at production scale.
- Obsessive focus on reliability and building systems that don't page you at 3am.
Nice-to-Haves
- Experience with Lance or other vector-native storage formats.
- Background in GPU-accelerated data processing (RAPIDS, cuDF).
Skills
Delta Lake, Iceberg, Hudi, Kafka, Flink, Ray, Spark, ClickHouse, Lance, Rapids, Cudf
Similar jobs
Data Engineering jobsBuild scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.
Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.