Staff Data Engineer
Staff Data Engineer leading data platform initiatives across batch, streaming, real-time pipelines, data lake infrastructure, governance, and privacy. Requires 5+ years of data engineering experience, strong Spark and distributed processing expertise, and the ability to lead complex production systems.
About the job
Responsibilities
- Design data solutions for batch, streaming, and real-time workloads.
- Design, build, maintain, and scale reliable data pipelines using Spark, Kafka, Iceberg, and Airflow/MWAA.
- Evolve the data lake and data platform across ingestion, processing, storage, and serving patterns.
- Build and operate cloud data infrastructure including Kafka, Airflow, Druid, and EMR on EKS, with strong reliability and observability.
- Implement data quality, observability, reliability, and governance across data pipelines and systems.
- Define standards for event instrumentation and governance, including event schemas, validation, and schema evolution.
- Build platform solutions for PII handling, data retention, deletion, and data residency supporting GDPR and CCPA requirements.
- Build AI agent harnesses that encode data engineering standards, pipeline patterns, and quality gates into development workflows.
- Lead complex technical initiatives from design through production, troubleshoot issues, mentor engineers, and collaborate with Analytics Engineering and other teams.
Requirements
- BA/BS degree or equivalent experience.
- 5+ years of experience in data engineering, building and operating production data pipelines and platforms.
- Strong experience with Spark and distributed data processing.
- Experience designing batch, real-time, and streaming solutions using technologies such as Kafka, Spark Structured Streaming, and CDC.
- Experience implementing data quality, observability, and reliability across production data systems.
- Experience with cloud data infrastructure, CI/CD, and infrastructure as code.
- Experience leading complex data engineering initiatives end to end.
- Strong programming and SQL skills.
Preferred Qualifications
- Experience with data governance and privacy, including PII management, retention, deletion, data residency, GDPR, or CCPA.
- Experience with event instrumentation and governance, including schemas, validation, and schema evolution.
- Experience evaluating architectural tradeoffs and establishing engineering best practices.
- Experience applying AI and emerging technologies to development workflows.
Compensation
- United States base salary ranges, depending on geographic zone:
- Zone A: $212,500–$255,000 USD
- Zone B: $200,000–$240,000 USD
- Zone C: $186,500–$224,000 USD
- Eligible for the company-wide bonus program and equity.
- Benefits include health coverage, paid parental leave, flexible vacation, holidays, sabbatical leave, wellness resources, a 401(k) with employer match, and monthly work and wellness stipends.
Skills
Spark, Apache Kafka, Apache Iceberg, Apache Airflow, Druid, Amazon Emr, Kubernetes, Spark Structured Streaming, SQL, Python, CI/CD, Infrastructure As Code, Data Governance, Data Quality, GDPR
Similar jobs
Data Engineering jobsBuild and scale data ingestion platforms, pipelines, APIs, and processing products that move billions of rows across a multi-tenant system. The role requires 8+ years of software development experience and strong expertise in large-scale application architecture.
Leads large-scale advertising data ingestion, measurement, and agentic workflow systems, combining deep ad-tech expertise with production LLM experience. Requires 10+ years of engineering experience and technical and people leadership in complex enterprise environments.
Staff-level engineer responsible for the technical direction, reliability, and evolution of a cloud ELT platform supporting healthcare data products. The role requires 7+ years of software or data engineering experience, deep SQL/Python and modern data-platform expertise, and strong architectural and mentoring leadership.
Leads organization-wide Snowflake migration, Medallion architecture, warehouse optimization, and CI/CD quality controls while partnering with executives on data strategy. The role requires 6+ years in analytics or data engineering, expert SQL, production dbt experience, and strong architectural judgment.
Leads enterprise data engineering strategy, architecture, delivery, governance, and technical leadership across the organization. Requires extensive data engineering experience, advanced data modeling and warehouse expertise, and strong PySpark, SQL, and Python skills.