Data Engineer
Build and operate scalable data pipelines, models, and infrastructure using Airflow, Snowflake, Databricks, AWS, and Terraform. The role requires 2+ years of data engineering experience, strong Python and SQL skills, and effective collaboration with technical and business stakeholders.
About the job
Responsibilities
- Build and maintain scalable data pipelines using Airflow to ingest, transform, and deliver data from various sources into Snowflake and Databricks.
- Design and implement Snowflake data models for analytics, reporting, and machine learning use cases, focusing on performance, reliability, and scalability.
- Develop infrastructure as code with Terraform to automate and manage AWS cloud resources.
- Monitor pipeline health and implement data quality checks for data accuracy, completeness, and timeliness.
- Optimize data processing workflows to improve performance, reduce costs, and handle growing data volumes.
- Troubleshoot pipeline issues, identify root causes, and implement long-term fixes.
- Collaborate with cross-functional teams across the US and India office to understand requirements and translate them into technical solutions.
- Document data pipelines, data models, and infrastructure, and facilitate knowledge transfer.
Requirements
- 2+ years of data engineering experience, with the ability to architect scalable data solutions.
- Hands-on experience with Python for data processing, automation, and data pipeline development.
- Proficiency with workflow orchestration tools, preferably Airflow, including DAG development, task dependencies, and monitoring.
- Strong SQL skills and experience with cloud data warehouses such as Snowflake, including performance optimization and data modeling.
- Experience with cloud platforms, preferably AWS, including S3, Lambda, EC2, and IAM.
- Experience collaborating with data analysts, analytics engineers, data scientists, and business stakeholders.
- Ownership of data pipeline reliability, performance, data flows, dependencies, and downstream implications.
- Proactive approach to independently driving work and structuring ambiguous technical problems.
Nice to Have
- Experience with dbt for transformation logic and analytics engineering workflows.
- Familiarity with Databricks, Spark optimization, and Delta Lake.
- Experience with infrastructure as code tools such as Terraform.
- Knowledge of dimensional modeling, star schemas, snowflake schemas, and slowly changing dimensions.
- Experience with CI/CD practices for data pipelines and automated testing frameworks.
- Knowledge of streaming data and real-time processing frameworks.
Skills
Python, Airflow, SQL, Snowflake, AWS, Amazon S3, AWS Lambda, Amazon Ec2, Aws Iam, Databricks, Terraform, dbt, Spark, Delta Lake, CI/CD
Similar jobs
Data Engineering jobsLeads the operational engine for collecting and annotating real-world and simulated data used by perception and robot-learning teams. The role manages vendors and annotators, quality systems, dataset governance, dashboards, and cross-functional delivery.
Build and operate secure batch and streaming data pipelines, dimensional models, and data quality systems for a healthcare data platform. Requires a bachelor’s degree and at least 3 years of data engineering experience, with strong Python, SQL, distributed systems, and data warehousing expertise.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Build and operate scalable lakehouse infrastructure, streaming and CDC pipelines, query systems, and self-serve BI capabilities. Requires 5+ years of data engineering experience, strong Kubernetes and infrastructure-as-code expertise, and hands-on experience with distributed data platforms.
Build and operate distributed systems powering Apache Pinot’s real-time analytics platform at massive scale. The role requires strong distributed-systems expertise, end-to-end delivery ownership, and a focus on reliability, observability, and performance.