Principal Data Engineer
Leads the strategy, architecture, and hands-on development of scalable AWS data ingestion and transformation platforms. Requires expert Python and SQL skills, Terraform and cloud-native pipeline experience, and 8+ years in data engineering or backend development, with technical leadership responsibilities.
About the job
Responsibilities
- Lead the technical strategy for the data platform and align infrastructure with business goals.
- Design and develop scalable, efficient, and reliable data ingestion and transformation pipelines across APIs, web data, and internal feeds.
- Provide technical leadership and mentorship to senior data platform engineers, including architecture, design, and implementation guidance.
- Architect and manage workflows using AWS MWAA (Airflow), Lambda, ECS, and SQS.
- Implement and maintain infrastructure as code with Terraform.
- Partner with data analysts, data scientists, and backend engineers to improve data consistency, discoverability, and reliability.
- Apply data modeling, schema design, and ETL/ELT best practices for high-volume structured and semi-structured data.
- Establish automated testing, monitoring, alerting, and lineage practices to ensure data quality.
- Promote code reviews, observability, continuous improvement, and knowledge sharing.
- Align data platform development with technology leadership, business strategy, and product goals.
Requirements
- Strong software engineering foundations, including SOLID design, modularity, and scalability.
- Expert Python proficiency for data pipelines and automation.
- Advanced SQL skills, including complex query and data model optimization.
- Experience designing and maintaining cloud-native AWS data pipelines using services such as MWAA/Airflow, Lambda, ECS, SQS, Glue, S3, and Redshift.
- Experience with data warehousing or lakehouse technologies such as Redshift, Snowflake, or Databricks.
- Experience managing Terraform or similar infrastructure-as-code frameworks.
- Strong understanding of data ingestion, transformation, orchestration, and AI/ML pipeline patterns.
- Familiarity with CI/CD, automated testing, and modern DevOps practices.
- 8+ years of experience in data engineering or backend development focused on scalable data solutions.
- Technical leadership experience, including mentoring and leading complex data infrastructure projects end to end.
- Familiarity with Docker and workflow orchestration best practices.
- Excellent communication, collaboration, and problem-solving skills.
Nice to Have
- Experience with streaming technologies such as Kafka, Kinesis, or Flink.
- Experience with data quality and observability tools such as Great Expectations or Monte Carlo.
- Familiarity with Scrapy, BeautifulSoup, or other data extraction frameworks.
Compensation and Benefits
- Health benefits.
- Matched 401(k) and pension plans.
- Paid time off.
- Generous parental leave.
- Gym subsidies.
- Educational reimbursement for career development.
- Recognition programs.
Skills
Python, SQL, AWS, Mwaa, Apache Airflow, AWS Lambda, Amazon Ecs, Amazon Sqs, Aws Glue, Amazon S3, Amazon Redshift, Terraform, Snowflake, Databricks, Docker
Similar jobs
Data Engineering jobsOwn the architecture, reliability, and evolution of a modern data platform spanning data engineering and analytics/BI. The role requires 8+ years of experience plus deep expertise in SQL, dbt, ClickHouse, BigQuery, Looker, streaming systems, and distributed data platforms.
Architects and builds ZoomInfo’s distributed data-platform infrastructure, including federated GraphQL access, real-time pipelines, indexing, and observability. The role requires 10+ years of software engineering experience, cloud-native expertise, and strong distributed-systems design skills.
Leads the strategy, architecture, and scaling of Pinterest’s big data and AI infrastructure across petabyte-scale workloads. Requires principal-level technical leadership, extensive Kubernetes or big data platform experience, and proficiency in modern data and cloud technologies.
Leads enterprise data engineering strategy, architecture, delivery, governance, and technical leadership across the organization. Requires extensive data engineering experience, advanced data modeling and warehouse expertise, and strong PySpark, SQL, and Python skills.
Staff Software Engineer leading development of data catalog and metadata infrastructure for discovery, governance, lineage, and quality across Airbnb’s data ecosystem. Requires 9+ years of software engineering experience focused on data infrastructure and strong programming and distributed data technology skills.