Skip to content
GranicaGranica

Senior Software Engineer — Lakehouse Systems

Build and optimize foundational lakehouse infrastructure for AI, spanning metadata, transactions, table maintenance, storage layout, and query performance at massive scale. The role requires senior systems engineering experience with modern lakehouse technologies, columnar formats, cloud object storage, and systems-oriented programming languages.

About the job

Responsibilities

  • Build foundational lakehouse systems for AI, including metadata management, transaction semantics, table maintenance, storage layouts, file-level optimization, and cost/performance optimization.
  • Design systems supporting time travel, schema evolution, partition evolution, snapshot isolation, and atomic consistency.
  • Develop infrastructure for manifests, snapshots, transaction logs, metadata pruning, snapshot expiration, table garbage collection, and catalog consistency.
  • Optimize file layout, clustering, compaction, file sizing, data skipping, indexing, and read-path performance.
  • Improve performance and cost efficiency across S3-, GCS-, and ADLS-backed lakehouse environments.
  • Optimize Parquet and ORC encoding, compression, layout, pruning, and read paths.
  • Build reliable, efficient lakehouse systems across Spark, Flink, Trino, Presto, Databricks, and related platforms.
  • Debug bottlenecks across storage, metadata, table maintenance, query execution, network, and compute layers.
  • Develop workload-aware optimization systems that learn from access patterns and automatically reorganize data.
  • Implement algorithms for compression, representation, layout optimization, and data efficiency.
  • Contribute to open source or publish research when appropriate.

Requirements

  • Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure.
  • Production experience with modern data lake or lakehouse technologies such as Apache Iceberg, Delta Lake, Apache Hudi, Spark, Trino, Presto, Flink, Hive Metastore, or Unity Catalog.
  • Hands-on experience with columnar formats such as Parquet or ORC.
  • Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout.
  • Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection.
  • Familiarity with cloud object storage such as S3, GCS, or ADLS and its performance tradeoffs.
  • Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages.
  • Curiosity about compression, entropy, information theory, and the effect of data representation on AI efficiency.
  • Pragmatic, rigorous, hands-on approach with the ability to own complex systems end to end.

Nice-to-haves

  • Contributions to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems.
  • Experience with manifests, snapshots, metadata catalogs, schema evolution, partition evolution, delete handling, transaction logs, or table garbage collection.
  • Experience addressing small-file problems, optimizing object-store access patterns, or improving table health at scale.
  • Background in storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization.
  • Research or open-source contributions in distributed systems, databases, storage, compression, indexing, or data processing.
  • Interest in how physical data representation affects model training, inference, retrieval, and reasoning efficiency.

Compensation & Benefits

  • Salary: $160,000–$240,000 annually.
  • Meaningful equity and performance bonus for top performers.
  • 401(k) with company match, comprehensive health coverage, unlimited PTO, catered meals, and support for research, publication, and conference participation.

Skills

Distributed Systems, Storage Systems, Databases, Apache Iceberg, Delta Lake, Apache Hudi, Spark, Parquet, Orc, Cloud Object Storage, Java, Scala, Go, Rust, C++

Underdog Fantasy

Underdog Fantasy

United States

Senior Analytics Engineer - Regulatory Reporting
$160k+/yrRemote5+ YOEData Engineering

Build and own the data platform powering regulatory reporting for a federally regulated prediction-markets business. The role combines dbt and SQL engineering, automated validation, root-cause investigations, audit support, and direct collaboration with regulators and compliance.

Vanta

Vanta

Remote

Senior Analytics Engineer
$157k+/yrRemote4+ YOEData Engineering

Senior Analytics Engineer responsible for designing complex data models, building scalable SQL pipelines, enabling AI tooling, and improving data infrastructure to support self-serve analytics, dashboards, and data science at Vanta. Requires 4+ years data experience, software engineering mindset, and expertise with modern analytics tools like dbt.

Deepgram

Deepgram

United States

AI Data Readiness Lead
$165k+/yrRemote5+ YOEData Engineering

Own the company’s metric governance program by defining canonical metrics, enforcing them in semantic and catalog systems, improving data quality, and validating AI-agent outputs. Requires 5+ years in analytics or analytics engineering, strong SQL, production semantic-layer ownership, and experience with AI evaluation and data governance.

Gusto

Gusto

Denver, CO
Senior Data Engineer
$155k+/yrHybrid10+ YOEData Engineering

The Senior Data Engineer will build scalable pipelines, ETL workflows, and data products while partnering with analytics, product, and engineering teams. The role requires 8–10+ years of experience, strong SQL and programming skills, cloud data-platform expertise, and a focus on reliability, quality, and automation.

tastytrade

tastytrade

Chicago, IL

Senior PostgreSQL Database Administrator / Database Engineer
$150k+/yrHybrid5+ YOEData Engineering

Own the performance, availability, backup and recovery, upgrades, and observability of high-throughput PostgreSQL environments across on-premises infrastructure and Amazon RDS. The senior role requires production Patroni failover experience, deep PostgreSQL expertise, Linux fundamentals, software development proficiency, and strong documentation skills.