Senior Data Engineer - Real World Data
Senior Data Engineer building scalable pipelines to transform EHR and claims data into analytics-ready assets while conducting hands-on real-world evidence analyses. Requires 5+ years experience with 2+ years in healthcare data, strong SQL/Python, Snowflake/dbt/Dagster, and familiarity with OMOP and causal inference frameworks.
About the job
Responsibilities
- Model and transform raw EHR and claims data into clean, canonical, and analytics-ready datasets using SQL, Python, and clinical standards like OMOP.
- Build and manage scalable data pipelines using Dagster for orchestration, dbt for transformation, and Snowflake as the primary compute and storage engine.
- Conduct hands-on RWD analyses to answer scientific and strategic research questions—including disease epidemiology, treatment patterns, patient journey characterization, and comparative effectiveness.
- Partner with Data Scientists and clinical leads to design and execute observational studies, translating scientific questions into well-structured, reproducible analyses.
- Implement data validation, completeness, and observability frameworks to ensure real-world datasets are accurate, comprehensive, and trustworthy for downstream research and product use.
- Apply Generative AI techniques within transformation and analysis layers to accelerate data structuring and insight generation.
- Communicate findings clearly to both technical and non-technical stakeholders, including summaries for portfolio teams and leadership.
Requirements
- 5+ years of experience in data engineering, with at least 2 years working in healthcare or life sciences, including direct exposure to EHR or claims datasets.
- Experience with ontologies and biomedical schemas (e.g. UMLS, LOINC, ICD9/10, MeSH) and understanding of modalities found within RWD — billing claims, lab results, visit notes.
- Fluency in SQL and Python; experience building and maintaining production-grade pipelines that support analytics or scientific workflows.
- Experience building longitudinal patient cohorts from EHR or claims data, including index date logic, washout periods, and follow-up window construction.
- Solid understanding of causal inference frameworks such as potential outcomes and target trial emulation.
- Working familiarity with real-world evidence study design concepts—such as active comparator new user designs, time-to-event outcomes, confounder adjustment, and causal discovery algorithms.
- Hands-on expertise with modern data infrastructure, such as Snowflake, dbt, and Dagster.
- Value clarity, documentation, and structured thinking—especially when working with complex healthcare data.
Nice-to-Haves
- Experience in regulated or privacy-sensitive data environments and familiarity with governance models for PHI or sensitive data.
- Prior experience working with commercial RWD vendors (e.g. Truveta, Optum, Komodo, IQVIA) and understanding the nuances of licensed claims and EHR datasets, including longitudinal patient journey construction and line-of-therapy sequencing.
Skills
SQL, Python, Snowflake, dbt, Dagster, Omop, Ehr, Claims Data, Umls, Loinc, Icd-9, Icd-10, Mesh, Causal Inference, Generative AI
Similar jobs
Data Engineering jobsSenior Data Infrastructure Engineer responsible for building and operating reliable, low-latency streaming and batch data systems that support AI products. Requires 5+ years of production data infrastructure experience and expertise with technologies such as Kafka, Flink, ClickHouse, and Terraform.
Build and operate petabyte-scale data infrastructure powering Discord’s insights and products. The role requires 5+ years of software engineering experience, strong programming skills, and experience with large-scale pipelines, streaming, orchestration, or data warehousing.
Build and lead the central data platform, covering ingestion, warehousing, orchestration, streaming, self-service frameworks, and trust layers. The role requires 5+ years of production data infrastructure experience, strong Python and SQL skills, and expertise with Snowflake and modern data tooling.
Build and operate large-scale revenue data pipelines powering billing and cost attribution, while improving reliability, latency, and correctness. The role requires strong Spark and Airflow experience, cross-functional problem-solving, and operational ownership of mission-critical production systems.
Build and operate low-latency systems that capture, normalize, and distribute real-time market data for institutional trading. The role requires backend engineering experience, Java or C++, market data infrastructure knowledge, and exchange connectivity expertise.