# Senior Software Engineer — Lakehouse Systems

**Company:** [Granica](https://hotfix.jobs/companies/granica)
**Location:** Mountain View, CA
**Role:** Data Engineering
**Salary:** $160k – $240k/yr
**Experience:** 5+ years
**Skills:** Distributed Systems, Storage Systems, Databases, Apache Iceberg, Delta Lake, Apache Hudi, Spark, Parquet, Orc, Cloud Object Storage, Java, Scala, Go, Rust, C++
**Posted:** 2026-08-27

> Build and optimize foundational lakehouse infrastructure for AI, spanning metadata, transactions, table maintenance, storage layout, and query performance at massive scale. The role requires senior systems engineering experience with modern lakehouse technologies, columnar formats, cloud object storage, and systems-oriented programming languages.

## Job Description

## Responsibilities
- Build foundational lakehouse systems for AI, including metadata management, transaction semantics, table maintenance, storage layouts, file-level optimization, and cost/performance optimization.
- Design systems supporting time travel, schema evolution, partition evolution, snapshot isolation, and atomic consistency.
- Develop infrastructure for manifests, snapshots, transaction logs, metadata pruning, snapshot expiration, table garbage collection, and catalog consistency.
- Optimize file layout, clustering, compaction, file sizing, data skipping, indexing, and read-path performance.
- Improve performance and cost efficiency across S3-, GCS-, and ADLS-backed lakehouse environments.
- Optimize Parquet and ORC encoding, compression, layout, pruning, and read paths.
- Build reliable, efficient lakehouse systems across Spark, Flink, Trino, Presto, Databricks, and related platforms.
- Debug bottlenecks across storage, metadata, table maintenance, query execution, network, and compute layers.
- Develop workload-aware optimization systems that learn from access patterns and automatically reorganize data.
- Implement algorithms for compression, representation, layout optimization, and data efficiency.
- Contribute to open source or publish research when appropriate.

## Requirements
- Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure.
- Production experience with modern data lake or lakehouse technologies such as Apache Iceberg, Delta Lake, Apache Hudi, Spark, Trino, Presto, Flink, Hive Metastore, or Unity Catalog.
- Hands-on experience with columnar formats such as Parquet or ORC.
- Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout.
- Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection.
- Familiarity with cloud object storage such as S3, GCS, or ADLS and its performance tradeoffs.
- Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages.
- Curiosity about compression, entropy, information theory, and the effect of data representation on AI efficiency.
- Pragmatic, rigorous, hands-on approach with the ability to own complex systems end to end.

## Nice-to-haves
- Contributions to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems.
- Experience with manifests, snapshots, metadata catalogs, schema evolution, partition evolution, delete handling, transaction logs, or table garbage collection.
- Experience addressing small-file problems, optimizing object-store access patterns, or improving table health at scale.
- Background in storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization.
- Research or open-source contributions in distributed systems, databases, storage, compression, indexing, or data processing.
- Interest in how physical data representation affects model training, inference, retrieval, and reasoning efficiency.

## Compensation & Benefits
- Salary: $160,000–$240,000 annually.
- Meaningful equity and performance bonus for top performers.
- 401(k) with company match, comprehensive health coverage, unlimited PTO, catered meals, and support for research, publication, and conference participation.

## Similar jobs

- [Senior Analytics Engineer - Regulatory Reporting](https://hotfix.jobs/jobs/5d52d4da-221d-43ef-be4f-b9d0abb1c09c) - Underdog Fantasy - Remote - $160k – $195k/yr
- [Senior Analytics Engineer](https://hotfix.jobs/jobs/35391cad-52d7-4eea-9bb7-e7fc82ebfb9b) - Vanta - Remote - $157k – $185k/yr
- [AI Data Readiness Lead](https://hotfix.jobs/jobs/456fe66b-aa42-4f78-a5f7-99231661da49) - Deepgram - Remote - $165k – $220k/yr
- [Senior Data Engineer](https://hotfix.jobs/jobs/8c19f913-ecb2-463a-bafe-2593664dc2a9) - Gusto - Denver, CO - $155k – $220k/yr
- [Senior PostgreSQL Database Administrator / Database Engineer](https://hotfix.jobs/jobs/88c11b80-8675-47c5-ad9b-29b2aa353a83) - tastytrade - Chicago, IL - $150k – $180k/yr

**Apply:** https://hotfix.jobs/jobs/eb6473ff-9bfd-4c33-9e59-07d9e92180e6
**Canonical:** https://hotfix.jobs/jobs/eb6473ff-9bfd-4c33-9e59-07d9e92180e6