Skip to content
StripeStripe

Staff Software Engineer, Datalake Platform

Leads architecture and operation of a petabyte-scale data lake platform, including Iceberg metastore services, object storage abstractions, migrations, authorization, and compliance controls. Requires 10+ years of software engineering experience and a track record with large-scale distributed storage or data infrastructure.

About the job

Responsibilities

  • Architect a unified Apache Iceberg platform, including a metastore service serving as the source of truth for Iceberg table management across Spark, Trino, Flink, and PyIceberg.
  • Define API contracts, authorization models, per-table credential vending, and integration patterns for data pipelines.
  • Own the migration strategy from Hive Metastore-backed workloads, including sequencing, backward compatibility, rollback, and cross-team coordination.
  • Define an object storage abstraction layer covering bucket provisioning, access-control policy design, and developer-facing client libraries.
  • Translate regulatory requirements into preventative technical controls, including audit logging, access review infrastructure, data segregation, and lifecycle enforcement.
  • Identify inefficiencies in storage layout, snapshot retention, and data lifecycle at petabyte scale; design automated self-service tooling to address them.
  • Lead critical design reviews, establish reliability, security, and developer-experience standards, and mentor senior engineers through architectural decisions.

Requirements

  • 10+ years of professional software engineering experience.
  • Track record designing, building, and operating large-scale distributed storage or data infrastructure systems.
  • Deep experience with object storage such as Amazon S3 or Azure Blob, including IAM, access-control policy design, lifecycle management, and petabyte-scale operations.
  • Experience leading complex, multi-quarter infrastructure projects end-to-end, managing cross-team dependencies, and coordinating migrations across many consuming teams.
  • Strong background in authorization and access-control design for distributed data systems.

Nice to Have

  • Deep expertise in Apache Iceberg, including table-format internals, the REST Catalog specification, snapshot lifecycle management, compaction, and integration with Spark, Trino, Flink, and PyIceberg.
  • Experience with compliance-sensitive infrastructure, including SOX, ICFR, or equivalent regulatory frameworks, and translating audit and access-review requirements into preventative technical controls.
  • Experience executing large-scale data migrations with attention to sequencing, blast-radius reduction, rollback, and data-integrity validation.
  • Strong developer-experience sensibility, including building ergonomic, well-documented abstractions that reduce engineering toil.

Skills

Apache Iceberg, Hive Metastore, Spark, Trino, Flink, Pyiceberg, Amazon S3, Azure Blob, IAM, Access Control, Authorization, Distributed Systems, Data Migration, Audit Logging, Lifecycle Management

Vanta

Vanta

Remote

Staff Software Engineer, Foundations
$260k+/yrRemote7+ YOEData Engineering

Leads the re-platforming of Vanta’s compliance data layer from MongoDB to schema-aware PostgreSQL across high-throughput Kafka and S3 pipelines. The role requires staff-level distributed systems expertise, migration leadership, and strong experience with relational and document data modeling.

Intercom

Intercom

Dublin, Ireland
Staff Data Engineer - GTM
No salary listedHybrid7+ YOEData Engineering

Builds the account, contact, hierarchy, enrichment, and identity systems that power go-to-market operations. The role requires modern data-stack experience, production LLM development, entity resolution, SaaS integrations, and close partnership with Sales, Marketing, and RevOps.

Alpaca

Alpaca

Remote

Senior Data Engineer
No salary listedRemote5+ YOEData Engineering

Build and operate scalable lakehouse infrastructure, streaming and CDC pipelines, query systems, and self-serve BI capabilities. Requires 5+ years of data engineering experience, strong Kubernetes and infrastructure-as-code expertise, and hands-on experience with distributed data platforms.

Thyme Care

Thyme Care

Remote

Senior Platform Engineer, Data
$176k+/yrRemote5+ YOEData Engineering

The Senior Platform Engineer will build and operate reliable data platform tooling, consolidate orchestration, scale dbt infrastructure, and improve Databricks developer experience. The role requires 5+ years of production software experience, strong Python and AWS expertise, infrastructure-as-code experience, and familiarity with modern data stacks.

Vanta

Vanta

Remote

Senior Data Engineer
$190k+/yrRemote4+ YOEData Engineering

Senior Data Engineer responsible for designing and deploying scalable data infrastructure, orchestration models, and analytics tooling to enable data-driven decisions, ML products, and enterprise reporting at Vanta. Requires 4+ years data experience, software engineering mindset, modern data stack proficiency, and passion for secure, compliant data systems.