Software Engineer - Data Platform
Builds and operates petabyte-scale data platform infrastructure using Kafka, Spark, Flink, and Trino to power real-time ML pipelines and analytics. Requires expertise in distributed systems, stream processing, and systems languages like Rust, Go, or Scala.
About the job
About the Team
The Data Platform team builds and operates infrastructure for large-scale data transport and processing, managing Apache Kafka, HDFS, Spark, Flink, and Trino for real-time ML pipelines, feed ranking, experimentation, analytics, and observability at petabyte scale.
About the Role
Design, build, and operate distributed systems powering data movement and compute, processing trillions of events daily for scalability, performance, and reliability in product and ML workloads.
What You Will Do
- Design and implement high-throughput, low-latency data ingestion and transport systems.
- Scale and optimize multi-tenant Kafka infrastructure supporting real-time workloads.
- Extend and tune Spark, Flink, and Trino for demanding production pipelines.
- Build interfaces, APIs, and pipelines enabling teams to query, process, and move data at petabyte scale.
- Debug and optimize distributed systems, with a focus on reliability and performance under load.
- Collaborate with ML, product, and infrastructure teams to unblock critical data workflows.
Ideal Candidate
- Proven expertise in distributed systems, stream processing, or large-scale data platforms.
- Proficiency in Rust, Go, Scala or similar systems languages.
- Hands-on experience with Kafka, Flink, Spark, Trino, or Hadoop in production.
- Strong debugging, profiling, and performance optimization skills.
- Track record of shipping and maintaining critical infrastructure.
- Comfortable working in fast-moving, high-stakes environments with minimal guardrails.
Compensation and Benefits
Annual Salary Range: $180,000 - $440,000 USD
Base salary plus equity, comprehensive medical, vision, dental, 401(k), disability insurance, life insurance, and perks.
Skills
Kafka, Spark, Flink, Trino, Hdfs, Rust, Go, Scala, Distributed Systems, Stream Processing
Similar jobs
Data Engineering jobsBuild scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.
Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.