Senior Data Engineer

Build and scale production data pipelines using Python, Spark, and Airflow to transform raw healthcare data into trusted datasets powering ML models, dashboards, and customer onboarding. Requires 6+ years experience with strong expertise in data engineering tools and cross-team collaboration.

180k – 220kUnited StatesData EngineeringRemote6+ YOE

Apply

About the role

What You’ll Do

Design and implement robust, production-grade pipelines using Python, Spark SQL, and Airflow to process high-volume file-based datasets (CSV, Parquet, JSON).
Lead efforts to canonicalize raw healthcare data (837 claims, EHR, partner data, flat files) into internal models.
Own the full lifecycle of core pipelines — from file ingestion to validated, queryable datasets — ensuring high reliability and performance.
Onboard new customers by integrating their raw data into internal pipelines and canonical models; collaborate with SMEs, Account Managers, and Product to ensure successful implementation and troubleshooting.
Build resilient, idempotent transformation logic with data quality checks, validation layers, and observability.
Refactor and scale existing pipelines to meet growing data and business needs.
Tune Spark jobs and optimize distributed processing performance.
Implement schema enforcement and versioning aligned with internal data standards.
Collaborate deeply with Data Analysts, Data Scientists, Product Managers, Engineering, Platform, SMEs, and AMs to ensure pipelines meet evolving business needs.
Monitor pipeline health, participate in on-call rotations, and proactively debug and resolve production data flow issues.
Contribute to the evolution of our data platform — driving toward mature patterns in observability, testing, and automation.
Build and enhance streaming pipelines (Kafka, SQS, or similar) where needed to support near-real-time data needs.
Help develop and champion internal best practices around pipeline development and data modeling.

What You Bring

6+ years of experience as a Data Engineer (or equivalent), building production-grade pipelines.
Strong expertise in Python, Spark SQL, and Airflow.
Experience processing large-scale file-based datasets (CSV, Parquet, JSON, etc) in production environments.
Experience mapping and standardizing raw external data into canonical models.
Familiarity with AWS (or any cloud), including file storage and distributed compute concepts.
Experience onboarding new customers and integrating external customer data with non-standard formats.
Ability to work across teams, manage priorities, and own complex data workflows with minimal supervision.
Strong written and verbal communication skills — able to explain technical concepts to non-engineering partners.
Comfortable designing pipelines from scratch and improving existing pipelines.
Experience working with large-scale or messy datasets (healthcare, financial, logs, etc.).
Experience building or willingness to learn streaming pipelines using tools such as Kafka or SQS.

Bonus: Familiarity with healthcare data (837, 835, EHR, UB04, claims normalization).

What We Offer

Work from anywhere in the US!
Full Medical/Dental/Vision for employees & their families.
Flexible and trusting environment.
Unlimited FTO.
Competitive salary, equity, 401(k) including employer match.
Base salary range: $180k-$220k.

Skills

PythonSpark SqlAirflowAWSKafkaSQSParquetJSONCsvSpark

Similar roles

Data Engineering jobs

Altana

Lead BI Engineer

Altana is seeking a founding Lead BI Engineer to build and own the company's analytical infrastructure. This role involves designing data warehouses, defining dimensional models, and creating executive dashboards to enable data-driven decision-making across the organization.

180k – 210kBrooklyn, NY +1Data EngineeringOn-site7+ YOESQLdbt

Garner Health

Senior Data Engineer

Builds, optimizes, and maintains scalable data pipelines using Python, SQL, and AWS tools for healthcare data insights. Requires 4+ years experience, expertise in data modeling, orchestration with Airflow, and familiarity with HIPAA.

180k – 220kNew York, NYData EngineeringHybrid4+ YOESQLAWS

Machinify

Senior Software Engineer (Data Platform)

Build scalable data platform backend systems using Golang/Python to support AI-driven healthcare solutions. Requires 6+ years backend experience, expertise in Kubernetes, cloud platforms, ETL, and data orchestration tools.

180k – 220kUnited StatesData EngineeringRemote6+ YOEGoAWS

Camber

Senior Data Engineer

Builds and scales reliable data pipelines and infrastructure using AWS and tools like Spark, Kafka, and dbt. Collaborates with teams to deliver data solutions, mentors engineers, and drives data architecture strategy. Requires 6+ years experience.

180k – 230kNew York, NYData EngineeringOn-site6+ YOEAWSdbt

Camber

Senior Software Engineer, Data

Build and scale data ingestion infrastructure and EHR integrations to power healthcare products, ensuring reliability for thousands of patients. Requires 5+ years in production systems, 3+ in data platforms, and expertise in Python, AWS, SQL, and big data tools like Spark and Kafka.

180k – 230kNew York, NYData EngineeringOn-site5+ YOEAWSSQL