Senior Data Engineer - Data Engineering
Builds and owns scalable SQL/Python data pipelines, golden datasets, and workflows using DBT, Airflow, Redshift for large-scale data (500TB+). Collaborates cross-functionally to enable data-driven decisions at Plaid. Requires 4+ years data engineering experience.
About the job
Responsibilities
- Understand different aspects of the Plaid product and strategy to inform golden dataset choices, design and data usage principles.
- Have data quality and performance top of mind while designing datasets.
- Lead key data engineering projects that drive collaboration across the company.
- Advocate for adopting industry tools and practices at the right time.
- Own core SQL and Python data pipelines that power our data lake and data warehouse.
- Deliver well-documented data with defined dataset quality, uptime, and usefulness.
Qualifications
- 4+ years of dedicated data engineering experience, solving complex data pipelines issues at scale.
- Experience building data models and data pipelines on top of large datasets (500TB to petabytes).
- Value SQL as a flexible tool, comfortable with modern SQL data orchestration tools like DBT, Mode, and Airflow.
- Experience with performant warehouses and data lakes: Redshift, Snowflake, Databricks.
- Experience building and maintaining batch and realtime pipelines using Spark, Kafka.
- Appreciation for schema design, evolving analytics schema on top of unstructured data.
- Excited to try new technologies, produce proof-of-concepts balancing technical advancement and adoption.
- Get deep into managing, deploying, and improving low-level data infrastructure.
- Empathetic with stakeholders, listen and collaborate on solutions balancing infra and business needs.
- Champion for data privacy and integrity.
Skills
SQL, Python, dbt, Airflow, Redshift, Snowflake, Databricks, Spark, Kafka, Data Pipelines
Similar jobs
Data Engineering jobsBuild and own large-scale data models, batch and real-time pipelines, and data infrastructure that provide reliable datasets and insights across Plaid. The role requires 4+ years of data engineering experience, strong SQL and Python skills, and expertise with modern warehouses, lakes, and orchestration tools.
Build and operate low-latency systems that capture, normalize, and distribute real-time market data for institutional trading. The role requires backend engineering experience, Java or C++, market data infrastructure knowledge, and exchange connectivity expertise.
Build and operate Jump’s Snowflake and dbt data platform, including ELT pipelines, governance, testing, monitoring, and analytics models. The role requires 6+ years of data engineering experience, production-grade SQL and Snowflake expertise, deep dbt knowledge, and strong Python and software engineering practices.
Senior Data Engineer responsible for designing and deploying scalable data infrastructure, orchestration models, and analytics tooling to enable data-driven decisions, ML products, and enterprise reporting at Vanta. Requires 4+ years data experience, software engineering mindset, modern data stack proficiency, and passion for secure, compliant data systems.
Build and operate large-scale revenue data pipelines powering billing and cost attribution, while improving reliability, latency, and correctness. The role requires strong Spark and Airflow experience, cross-functional problem-solving, and operational ownership of mission-critical production systems.