Leads architecture and evolution of a large-scale data platform spanning ingestion, lakehouse storage, streaming, governance, and access. Requires 10+ years of software engineering experience, deep production expertise with Spark and distributed systems, and strong cloud data-platform experience.
Salary not listed
Hybrid7+ YOEData Engineering
About the role
Responsibilities
Lead the design of scalable and reliable data pipelines, storage solutions, and access layers using modern cloud-native technologies.
Own and evolve core data platform components, including event streaming, lakehouse architecture, batch and streaming ETL, and data cataloging.
Define best practices for data quality, lineage, privacy, and access control to ensure regulatory compliance and trust in the data.
Collaborate with Product, Analytics, Infrastructure, and Security teams to align data platform capabilities with organizational goals.
Mentor engineers across the organization, establish coding and architectural standards, and influence the strategic direction of the platform.
Evaluate and integrate technologies that improve performance, observability, developer experience, and cost efficiency.
Define compute architecture and standards for Spark jobs across the data platform.
Build and operate production-grade Spark pipelines using Apache Iceberg, Kafka, and cloud-native data services.
Requirements
10+ years of software engineering experience with deep expertise in data infrastructure, distributed systems, or backend platform engineering.
Proven experience designing and delivering production-grade data platforms at scale, including platforms supporting hundreds of terabytes of data and thousands of jobs per day.
Strong experience with Snowflake, dbt, Kafka, Airflow, Spark, and cloud-native technologies such as AWS, GCP, or Azure.
Hands-on experience building APIs and services for data access, metadata management, and platform observability.
Familiarity with GDPR, SOC 2, HIPAA, and enterprise-grade security practices.
Effective communication and collaboration skills, with the ability to influence engineering and non-engineering audiences.
Experience in a startup or high-growth SaaS environment.
Nice-to-Haves
Exposure to AI/ML data pipelines or real-time analytics platforms.
Production experience with Apache Iceberg, Structured Streaming, Kafka on MSK, AWS S3, AWS Glue, MWAA, Trino, or Athena.
Expertise in Spark performance tuning, including shuffle behavior, partition strategies, memory pressure, Catalyst optimization, and Spark UI analysis.
Experience choosing between PySpark and Scala Spark for production workloads.
Compensation and Benefits
The posting does not specify compensation or benefits.
Builds and guides development of backend business-intelligence tooling, data pipelines, and high-availability platform services. The role requires a bachelor's degree, 7 years of backend and data-platform experience, and expertise across Java, Python, SQL, AWS, distributed systems, and streaming technologies.
193k – 240k/yrRemote7+ YOEData Engineering
Staff Data Engineer - Data Infrastructure
DiscordUnited States
Leads the technical strategy, governance, and quality of Discord’s analytical data infrastructure, building trusted curated datasets and metric frameworks for internal and externally reported analytics. Requires 7+ years in data and software engineering, expert SQL, strong Python, and experience operating data systems at scale.
279k – 310k/yrOn-site7+ YOEData Engineering
Staff Software Engineer, Data Governance & Foundations
InstacartUnited States
Leads architecture and delivery of Instacart’s open lakehouse foundation, governance controls, and multi-engine compute strategy. Requires 10+ years building production-scale data infrastructure or distributed systems, with expertise in lakehouse, streaming, and platform migrations.
221k – 280k/yrRemote10+ YOEData Engineering
Staff Analytics Engineer, Compliance Data
CoinbaseUnited States
Leads the architecture, modeling, pipelines, quality systems, and certification processes for a canonical compliance data platform. Requires 8+ years in analytics or data engineering, expert SQL and Python, modern warehouse expertise, and technical leadership under regulatory pressure.
207k – 244k/yrRemote8+ YOEData Engineering
Staff Data Engineer
CriblCalifornia
Leads architecture and operations for Cribl’s data platform, including Snowflake, Prefect, dbt, AWS infrastructure, ingestion services, and AI-ready data workflows. The role requires deep Snowflake expertise, strong Python and SQL skills, infrastructure-as-code experience, and the ability to mentor engineers and set technical direction.