Senior Software Engineer, Data Platform
Build and lead the central data platform, covering ingestion, warehousing, orchestration, streaming, self-service frameworks, and trust layers. The role requires 5+ years of production data infrastructure experience, strong Python and SQL skills, and expertise with Snowflake and modern data tooling.
About the job
Responsibilities
- Own the data platform architecture and technical direction, including reusable frameworks and build-versus-buy decisions.
- Build and operate ingestion across streaming, batch, CDC, and third-party connectors, with safe schema evolution.
- Land data in Snowflake with reliable freshness, completeness, and cost characteristics.
- Own orchestration for scheduling, retries, backfills, and dependency management.
- Build transformation, compute, and self-service pipeline frameworks.
- Design and operate stream-processing infrastructure for real-time products, alerting, and reporting.
- Build data-quality, observability, lineage, cataloging, and discovery capabilities.
- Develop tooling for PII classification, masking, retention, access control, and multi-region data residency.
- Set technical standards through design reviews, documentation, and mentorship.
Requirements
- 5+ years building and operating production data infrastructure.
- Deep experience with cloud data warehouses; Snowflake is strongly preferred.
- Experience with CDC and streaming pipelines using technologies such as Kafka, Debezium, Flink, or Spark Streaming.
- Experience with managed ingestion tools such as Fivetran or Airbyte.
- Strong fluency with workflow orchestration tools such as Temporal, Airflow, or Dagster.
- Strong programming skills in Python and advanced SQL.
- Experience building frameworks or internal tooling for other engineers.
- Experience with data quality, observability, lineage, and schema evolution.
- Working knowledge of data governance in regulated environments, including PII classification, masking, access control, retention, and data residency.
- Familiarity with Azure, AWS, or GCP; Kubernetes; and infrastructure-as-code tools such as Terraform or Pulumi.
- Comfort operating in ambiguity and defining scope.
Nice to Have
- Experience with dbt and analytics engineering teams.
- Experience with lakehouse architectures, Iceberg, Delta Lake, or Trino.
- Experience operating multi-tenant platforms with strict security, compliance, or data residency requirements.
- Exposure to data infrastructure for AI products.
- Experience as an early or founding data platform hire at a fast-growing company.
Compensation
- $193,400–$290,000 USD
Skills
Snowflake, Kafka, Debezium, Flink, Spark, Fivetran, Airbyte, Temporal, Apache Airflow, Dagster, Python, SQL, dbt, Kubernetes, Terraform
Similar jobs
Data Engineering jobsBuild and operate large-scale revenue data pipelines powering billing and cost attribution, while improving reliability, latency, and correctness. The role requires strong Spark and Airflow experience, cross-functional problem-solving, and operational ownership of mission-critical production systems.
Build and operate low-latency systems that capture, normalize, and distribute real-time market data for institutional trading. The role requires backend engineering experience, Java or C++, market data infrastructure knowledge, and exchange connectivity expertise.
Build and own large-scale data models, batch and real-time pipelines, and data infrastructure that provide reliable datasets and insights across Plaid. The role requires 4+ years of data engineering experience, strong SQL and Python skills, and expertise with modern warehouses, lakes, and orchestration tools.
Builds and owns scalable SQL/Python data pipelines, golden datasets, and workflows using DBT, Airflow, Redshift for large-scale data (500TB+). Collaborates cross-functionally to enable data-driven decisions at Plaid. Requires 4+ years data engineering experience.
Build and operate petabyte-scale data infrastructure powering Discord’s insights and products. The role requires 5+ years of software engineering experience, strong programming skills, and experience with large-scale pipelines, streaming, orchestration, or data warehousing.