Senior Data Engineer
Build and scale production data pipelines that transform raw healthcare and customer data into trusted canonical datasets powering ML models, dashboards, and product decisions. The role requires 6+ years of data engineering experience and strong Python, Spark SQL, and Airflow expertise.
About the job
Responsibilities
- Design and implement production-grade pipelines using Python, Spark SQL, and Airflow for high-volume CSV, Parquet, and JSON datasets.
- Canonicalize raw healthcare data, including 837 claims, EHR, partner data, and flat files, into internal models.
- Own pipelines from file ingestion through validated, queryable datasets, ensuring reliability and performance.
- Onboard customers by integrating external data into internal pipelines and canonical models; collaborate with subject-matter experts, account managers, and Product on implementation and troubleshooting.
- Build resilient, idempotent transformation logic with data-quality checks, validation layers, schema enforcement, versioning, and observability.
- Refactor and scale pipelines, tune Spark jobs, and optimize distributed processing.
- Collaborate with Data Analysts, Data Scientists, Product Managers, Engineering, Platform, subject-matter experts, and account managers.
- Monitor pipeline health, participate in on-call rotations, and debug production data-flow issues.
- Contribute to data-platform patterns for observability, testing, automation, data modeling, and pipeline development.
- Build or enhance streaming pipelines using Kafka, SQS, or similar tools.
Requirements
- 6+ years of experience as a Data Engineer or equivalent building production-grade pipelines.
- Strong expertise in Python, Spark SQL, and Airflow.
- Experience processing large-scale file-based datasets such as CSV, Parquet, and JSON in production.
- Experience mapping and standardizing raw external data into canonical models.
- Familiarity with AWS or another cloud platform, including file storage and distributed-compute concepts.
- Experience onboarding customers and integrating external data with non-standard formats.
- Ability to manage priorities, work across teams, and own complex data workflows with minimal supervision.
- Strong written and verbal communication skills.
- Experience designing pipelines from scratch and improving existing pipelines.
- Experience with large-scale or messy datasets.
- Experience building, or willingness to learn, streaming pipelines with Kafka or SQS.
Nice-to-haves
- Familiarity with healthcare data, including 837, 835, EHR, UB04, and claims normalization.
Compensation and Benefits
- Base salary range: $180,000-$220,000 annually.
- Equity and 401(k) with employer match.
- Medical, dental, and vision coverage for employees and families.
- Unlimited flexible time off.
Skills
Python, Spark Sql, Apache Airflow, AWS, Apache Kafka, Amazon Sqs, Spark, Csv, Parquet, JSON, Data Modeling, Data Quality
Similar jobs
Data Engineering jobsSenior Data Engineer responsible for designing and operating scalable data pipelines and platform capabilities across Snowflake and AWS. The role requires 5+ years of production data engineering experience, strong SQL and Python skills, and expertise in ETL/ELT, orchestration, quality, and observability.
Senior software engineer building and evolving Fetch’s data platform, including pipelines, governed data access, delivery infrastructure, and partner integrations. The role requires 8+ years of experience, strong platform or backend expertise, ownership of complex cross-team initiatives, and excellent technical judgment.
The Senior Platform Engineer will build and operate reliable data platform tooling, consolidate orchestration, scale dbt infrastructure, and improve Databricks developer experience. The role requires 5+ years of production software experience, strong Python and AWS expertise, infrastructure-as-code experience, and familiarity with modern data stacks.
The Senior Data Engineer will design scalable data pipelines and warehousing systems supporting analytics, business metrics, and machine-learning initiatives. The role requires 4+ years of enterprise data experience, expertise with modern data platforms and ETL, and the ability to mentor engineers and collaborate across functions.
Build and own foundational data models, reliable pipelines, and data platforms that support analytics and business decision-making. The role requires 5+ years of data engineering experience, strong Python and SQL skills, cloud data warehouse expertise, and a bachelor's degree or equivalent experience.