Skip to content

Senior Data Engineer Backend

Senior Data Engineer responsible for building and operating reliable clinical and claims data pipelines, CDC systems, quality controls, and de-identified exports. The role requires 5+ years of production pipeline experience plus strong SQL, Python, Spark, and data-governance skills.

About the job

Responsibilities

  • Build and operate ingestion pipelines from EHR systems into Databricks, including landing, validation, normalization, and silver and gold tables.
  • Onboard new clinics by understanding export formats and implementing scheduled orchestration.
  • Maintain change-data-capture from the MongoDB application database into the analytics layer and ensure consistency.
  • Build de-identified data exports for research partners with controls preventing identifiers from leaving the platform.
  • Own data quality through freshness checks, source reconciliation, and proactive alerting.
  • Investigate data incidents, identify root causes, and perform safe backfills.
  • Write orchestration code in TypeScript alongside the backend team.

Requirements

  • 5+ years of experience building and operating production data pipelines.
  • Strong production experience with SQL and Python.
  • Hands-on experience with Spark or a comparable engine and Delta Lake or an equivalent table format.
  • Knowledge of idempotency, deduplication keys, event ordering, and schema evolution.
  • Experience with change-data-capture or event-stream processing from an operational database.
  • Willingness to write TypeScript for orchestration code.
  • Understanding of sensitive-data leakage risks, including through file names, object keys, and logs.
  • Comfort solving ambiguous data-integration and data-quality problems.

Nice to Have

  • Familiarity with Databricks and Unity Catalog, including governance and access controls.
  • Experience with Temporal or another durable-execution engine.
  • Hands-on experience with MongoDB.
  • Familiarity with healthcare data, including claims, CPT and ICD-10 codes, and HIPAA de-identification rules.
  • Prior experience with TypeScript or Node.js.
  • Experience with infrastructure as code using Terraform and CI/CD for data pipelines.

Skills

SQL, Python, Spark, Delta Lake, Databricks, TypeScript, MongoDB, Change Data Capture, Temporal, Terraform, CI/CD, Unity Catalog, HIPAA

Alpaca

Alpaca

Remote

Senior Data Engineer
No salary listedRemote5+ YOEData Engineering

Build and operate scalable lakehouse infrastructure, streaming and CDC pipelines, query systems, and self-serve BI capabilities. Requires 5+ years of data engineering experience, strong Kubernetes and infrastructure-as-code expertise, and hands-on experience with distributed data platforms.

Lyft

Lyft

Toronto, Canada

Senior Software Engineer, Data - Mapping
CA$136k+/yrHybrid5+ YOEData Engineering

Leads architecture for offline experimentation and route simulation while building reliable, scalable data pipelines and backend services. The role requires 5+ years of backend or data engineering experience, distributed-systems expertise, strong SQL and Spark skills, and proficiency with modern cloud infrastructure.

Twilio

Twilio

United States
Senior Data and AI Specialist
$106k+/yrRemote5+ YOEData Engineering

Builds agentic AI, automated data workflows, and BI solutions for complex telecommunications datasets. The role requires 5+ years of technical data and automation experience, strong SQL and Python skills, and expertise in data governance and LLM-based tools.

Twilio

Twilio

Ontario, Canada
Senior Data and AI Specialist
CA$100k+/yrRemote5+ YOEData Engineering

The Senior Data and AI Specialist will build agentic AI solutions, automated Python workflows, and analytics products across telecommunications data. The role requires at least five years of experience with SQL, data automation, BI tools, LLM agents, and data governance.

Wrapbook

Wrapbook

United States
Senior Analytics Engineer II
CA$148k+/yrRemote5+ YOEData Engineering

Build and own Wrapbook’s analytics layer, including production pipelines, governed data models, canonical datasets, self-serve analytics, and monitoring. The role requires strong SQL and Python, modern warehouse experience, and 4+ years in data or analytics engineering.