Senior Data Engineer
Build and own large-scale data models, batch and real-time pipelines, and data infrastructure that provide reliable datasets and insights across Plaid. The role requires 4+ years of data engineering experience, strong SQL and Python skills, and expertise with modern warehouses, lakes, and orchestration tools.
About the job
Responsibilities
- Understand Plaid’s products and strategy to inform golden dataset selection, design, and data usage principles.
- Design datasets with data quality and performance in mind.
- Lead data engineering projects that drive collaboration across the company.
- Advocate for appropriate industry tools and practices.
- Own core SQL and Python data pipelines powering the data lake and data warehouse.
- Document data and define dataset quality, uptime, and usefulness SLAs.
- Collaborate with Engineering, Product, Marketing, Finance, business intelligence, data analysts, and other stakeholders.
Requirements
- 4+ years of dedicated data engineering experience solving complex data pipeline issues at scale.
- Experience building data models and pipelines on large datasets, approximately 500 TB to petabytes.
- Strong SQL skills and experience with modern SQL orchestration tools such as DBT, Mode, and Airflow.
- Experience with performant data warehouses and data lakes, including Redshift, Snowflake, or Databricks.
- Experience building and maintaining batch and real-time pipelines with technologies such as Spark and Kafka.
- Understanding of schema design and evolving analytics schemas over unstructured data.
- Ability to manage, deploy, and improve low-level data infrastructure.
- Strong stakeholder collaboration and problem-solving skills.
- Commitment to data privacy and integrity.
Compensation and Benefits
- Annual salary range: $190,800–$238,800.
- Additional compensation may include equity and/or commission, depending on the position.
- Comprehensive benefits include medical, dental, vision, and 401(k).
Skills
SQL, Python, dbt, Airflow, Amazon Redshift, Snowflake, Databricks, Spark, Apache Kafka, Data Modeling, Data Pipelines, Schema Design, Data Warehousing, Data Lakes, Real-Time Pipelines
Similar jobs
Data Engineering jobsBuilds and owns scalable SQL/Python data pipelines, golden datasets, and workflows using DBT, Airflow, Redshift for large-scale data (500TB+). Collaborates cross-functionally to enable data-driven decisions at Plaid. Requires 4+ years data engineering experience.
Build and operate low-latency systems that capture, normalize, and distribute real-time market data for institutional trading. The role requires backend engineering experience, Java or C++, market data infrastructure knowledge, and exchange connectivity expertise.
Build and operate Jump’s Snowflake and dbt data platform, including ELT pipelines, governance, testing, monitoring, and analytics models. The role requires 6+ years of data engineering experience, production-grade SQL and Snowflake expertise, deep dbt knowledge, and strong Python and software engineering practices.
Senior Data Engineer responsible for designing and deploying scalable data infrastructure, orchestration models, and analytics tooling to enable data-driven decisions, ML products, and enterprise reporting at Vanta. Requires 4+ years data experience, software engineering mindset, modern data stack proficiency, and passion for secure, compliant data systems.
Build and operate large-scale revenue data pipelines powering billing and cost attribution, while improving reliability, latency, and correctness. The role requires strong Spark and Airflow experience, cross-functional problem-solving, and operational ownership of mission-critical production systems.