Senior Data Engineer building and scaling clinical data pipelines, lakehouses, warehouses, and ETL/ELT systems for millions of patients to power real-time healthcare products and AI/ML workflows. Requires 6+ years data engineering experience, strong software engineering fundamentals, and mentoring skills.
180k – 220k/yr
Hybrid6+ YOEData Engineering
About the role
Responsibilities
Raise the technical bar for our data platform: building on and improving our warehouse, data lake, and ETL/ELT architecture so it scales with patient and customer growth, and helping evaluate the right tools (batch and streaming processing, table formats, orchestration, query engines).
Drive data projects end-to-end: writing Design Documents, shipping v0's quickly, and iterating to v1 and beyond.
Support AI/ML efforts - making sure the AI Engineers have the data they need.
Multiply the team: mentoring engineers on data fundamentals, reviewing designs and PRs, and judging when to invest in quality vs. ship fast.
Participate in bi-weekly sprint planning and retros, joining our daily 30-min remote stand-up at 7:30am PST (our only mandatory meeting), and taking part in the on-call rotation.
Example projects:
Scaling our patient data consolidation pipeline (deduplication, normalization, hydration) to handle 100x today's volume without 100x the cost.
Building pipelines that deliver clinical data directly into customers' data warehouses, reliably and at scale.
Building the ingestion path for customers pushing large volumes of their own data into the platform.
Building document-processing pipelines that extract structured data from PDFs, images, and free text to feed ML models.
Requirements
6+ years of engineering experience, with a heavy lean towards data engineering — building, maintaining, and scaling pipelines processing terabytes of data and millions of events a day.
Experience across the data stack — ingestion, storage, processing, warehousing, serving — and an understanding of the tradeoffs (cost, latency, correctness, operability) at each layer.
Experience with modern, cloud-native data stacks: e.g., Spark, open table formats (Parquet, Iceberg, Delta) on S3, and warehouses (Snowflake, BigQuery, Redshift).
Strong software engineering fundamentals — you write production code (we're a TypeScript shop, with Python in data/ML workflows), not just orchestration configs.
Experience mentoring or guiding other engineers — through code reviews, pairing, design feedback, or onboarding.
Located in San Francisco / Bay Area, or willing to relocate.
Nice-to-Haves
Experience with streaming systems (Kafka, Kinesis), dbt, or orchestration tooling (Airflow, Dagster).
Experience building or supporting ML/data science workflows (feature pipelines, model inputs/outputs, unstructured data extraction).
Lead data operations for AI audio models at HappyRobot. Translate research needs into data specs, manage end-to-end acquisition/annotation with internal and vendor teams, optimize tooling/workflows, and scale labeling operations to deliver high-quality datasets on time.
180k – 240k/yrHybrid5+ YOEData Engineering
Senior Data Engineer
Garner HealthNew York, NY
Builds, optimizes, and maintains scalable data pipelines using Python, SQL, and AWS tools for healthcare data insights. Requires 4+ years experience, expertise in data modeling, orchestration with Airflow, and familiarity with HIPAA.
180k – 220k/yrHybrid4+ YOEData Engineering
Senior Software Engineer (Data Platform)
MachinifyUnited States
Build scalable data platform backend systems using Golang/Python to support AI-driven healthcare solutions. Requires 6+ years backend experience, expertise in Kubernetes, cloud platforms, ETL, and data orchestration tools.
180k – 220k/yrRemote6+ YOEData Engineering
Senior Data Engineer
CamberNew York, NY
Builds and scales reliable data pipelines and infrastructure using AWS and tools like Spark, Kafka, and dbt. Collaborates with teams to deliver data solutions, mentors engineers, and drives data architecture strategy. Requires 6+ years experience.
180k – 230k/yrOn-site6+ YOEData Engineering
Senior Software Engineer, Data
CamberNew York, NY
Build and scale data ingestion infrastructure and EHR integrations to power healthcare products, ensuring reliability for thousands of patients. Requires 5+ years in production systems, 3+ in data platforms, and expertise in Python, AWS, SQL, and big data tools like Spark and Kafka.