Data Engineer
Build and operate scalable data pipelines, models, and infrastructure using Airflow, Snowflake, Databricks, AWS, and Terraform. The role requires 2+ years of data engineering experience, strong Python and SQL skills, and effective collaboration with technical and business stakeholders.
About the job
Responsibilities
- Build and maintain scalable data pipelines using Airflow to ingest, transform, and deliver data from various sources into Snowflake and Databricks.
- Design and implement Snowflake data models for analytics, reporting, and machine learning use cases, focusing on performance, reliability, and scalability.
- Develop infrastructure as code with Terraform to automate and manage AWS cloud resources.
- Monitor pipeline health and implement data quality checks for data accuracy, completeness, and timeliness.
- Optimize data processing workflows to improve performance, reduce costs, and handle growing data volumes.
- Troubleshoot pipeline issues, identify root causes, and implement long-term fixes.
- Collaborate with cross-functional teams across the US and India office to understand requirements and translate them into technical solutions.
- Document data pipelines, data models, and infrastructure, and facilitate knowledge transfer.
Requirements
- 2+ years of data engineering experience, with the ability to architect scalable data solutions.
- Hands-on experience with Python for data processing, automation, and data pipeline development.
- Proficiency with workflow orchestration tools, preferably Airflow, including DAG development, task dependencies, and monitoring.
- Strong SQL skills and experience with cloud data warehouses such as Snowflake, including performance optimization and data modeling.
- Experience with cloud platforms, preferably AWS, including S3, Lambda, EC2, and IAM.
- Experience collaborating with data analysts, analytics engineers, data scientists, and business stakeholders.
- Ownership of data pipeline reliability, performance, data flows, dependencies, and downstream implications.
- Proactive approach to independently driving work and structuring ambiguous technical problems.
Nice to Have
- Experience with dbt for transformation logic and analytics engineering workflows.
- Familiarity with Databricks, Spark optimization, and Delta Lake.
- Experience with infrastructure as code tools such as Terraform.
- Knowledge of dimensional modeling, star schemas, snowflake schemas, and slowly changing dimensions.
- Experience with CI/CD practices for data pipelines and automated testing frameworks.
- Knowledge of streaming data and real-time processing frameworks.
Skills
Python, Airflow, SQL, Snowflake, AWS, Amazon S3, AWS Lambda, Amazon Ec2, Aws Iam, Databricks, Terraform, dbt, Spark, Delta Lake, CI/CD
Similar jobs
Data Engineering jobsLeads the operational engine for collecting and annotating real-world and simulated data used by perception and robot-learning teams. The role manages vendors and annotators, quality systems, dataset governance, dashboards, and cross-functional delivery.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Build and operate scalable lakehouse infrastructure, streaming and CDC pipelines, query systems, and self-serve BI capabilities. Requires 5+ years of data engineering experience, strong Kubernetes and infrastructure-as-code expertise, and hands-on experience with distributed data platforms.
Build and operate distributed systems powering Apache Pinot’s real-time analytics platform at massive scale. The role requires strong distributed-systems expertise, end-to-end delivery ownership, and a focus on reliability, observability, and performance.
Senior data platform engineer who scales infrastructure, automates data delivery, builds AI-enabled analytical tools, and leads cross-functional engineering initiatives. Requires 4+ years of data infrastructure experience, strong Kafka and distributed-systems expertise, and proficiency in Python, Scala, cloud platforms, and Terraform.