Senior Software Engineer - Distributed Data Systems
Build distributed data storage and processing systems for big data workloads including Apache Spark, Delta Lake, and performance optimization. Requires 5+ years in Java/Scala/C++, strong algorithms knowledge, and distributed systems experience.
About the job
Responsibilities
- Build next generation distributed data storage and processing systems that outperform SQL query engines while supporting ETL, data science, and other workloads.
- Example projects:
- Apache Spark™: Develop the open source standard framework for big data.
- Data Plane Storage: Build reliable, high-performance services and client libraries for cloud storage (AWS S3, Azure Blob Store).
- Delta Lake: Create storage management system with ACID transactions, time travel, data lake scale, and warehouse performance.
- Delta Pipelines: Orchestrate and operate tens of thousands of data pipelines with higher-level abstractions.
- Performance Engineering: Develop fast, scalable query optimizer and execution engine.
Requirements
- BS (or higher) in Computer Science or related field, or equivalent experience.
- 5+ years production experience in Java, Scala, or C++.
- Strong foundation in algorithms, data structures, and real-world applications.
- Experience with distributed systems, databases, and big data systems (Apache Spark, Hadoop).
- Comfortable working toward multi-year vision with incremental deliverables.
- Motivated by customer value and impact.
Skills
Java, Scala, C++, Spark, Delta Lake, Aws S3, Azure Blob Storage, Algorithms, Data Structures, Distributed Systems
Similar jobs
Data Engineering jobsSenior Analytics Engineer responsible for designing complex data models, building scalable SQL pipelines, enabling AI tooling, and improving data infrastructure to support self-serve analytics, dashboards, and data science at Vanta. Requires 4+ years data experience, software engineering mindset, and expertise with modern analytics tools like dbt.
Build and own the data platform powering regulatory reporting for a federally regulated prediction-markets business. The role combines dbt and SQL engineering, automated validation, root-cause investigations, audit support, and direct collaboration with regulators and compliance.
Build and optimize foundational lakehouse infrastructure for AI, spanning metadata, transactions, table maintenance, storage layout, and query performance at massive scale. The role requires senior systems engineering experience with modern lakehouse technologies, columnar formats, cloud object storage, and systems-oriented programming languages.
The Senior Data Engineer will build scalable pipelines, ETL workflows, and data products while partnering with analytics, product, and engineering teams. The role requires 8–10+ years of experience, strong SQL and programming skills, cloud data-platform expertise, and a focus on reliability, quality, and automation.
Own the company’s metric governance program by defining canonical metrics, enforcing them in semantic and catalog systems, improving data quality, and validating AI-agent outputs. Requires 5+ years in analytics or analytics engineering, strong SQL, production semantic-layer ownership, and experience with AI evaluation and data governance.