Staff Data Engineer
Lead the design, architecture, and scaling of Metriport's data platform for ingesting and processing real-time clinical data from millions of patients. Own end-to-end data projects, mentor engineers, support AI/ML workflows, and eventually lead a team while staying hands-on. Requires 8+ years building large-scale data platforms with modern cloud-native tools.
About the job
What you'll be doing
We ingest clinical data for millions of patients from external healthcare sources, with continuous updates for a growing subset of those patients. You'll own the architecture and evolution of the data platform that powers our product — and ship it to customers fast.
Day to day, that looks like:
- Setting the technical direction for our data platform: evolving our warehouse, data lake, and ETL/ELT architecture to scale with patient and customer growth, and picking the right tools (batch and streaming processing, table formats, orchestration, query engines).
- Driving the critical data projects end-to-end: writing Design Documents, shipping v0's quickly, and iterating to v1 and beyond.
- Supporting AI/ML efforts: making sure the AI Engineers have the data they need.
- Multiplying the team: mentoring engineers on data fundamentals, reviewing designs and PRs, and judging when to invest in quality vs. ship fast.
- Eventually, acting as Team Lead for a group of engineers: breaking down and delegating work, unblocking teammates, and owning your team's delivery — while staying hands-on in the code.
- Driving bi-weekly sprint planning and retros, contributing to the engineering roadmap, joining our daily 30-min remote stand-up at 7:30am PST (our only mandatory meeting), and taking part in the on-call rotation.
Example projects you could own:
- Rearchitecting our patient data consolidation pipeline (deduplication, normalization, hydration) to handle 100x today's volume without 100x the cost.
- Building pipelines that deliver clinical data directly into customers' data warehouses, reliably and at scale.
- Designing the ingestion path for customers pushing large volumes of their own data into the platform.
- Building document-processing pipelines that extract structured data from PDFs, images, and free text to feed ML models.
Requirements
- 8+ years of engineering experience, with significant depth building, operating, and scaling data platforms processing terabytes of data and millions-to-billions of events a day.
- You've designed data architectures end-to-end — ingestion, storage, processing, warehousing, serving — and owned the tradeoffs (cost, latency, correctness, operability) at each layer.
- Deep experience with modern, cloud-native data stacks: e.g., Spark, open table formats (Parquet, Iceberg, Delta) on S3, warehouses (Snowflake, BigQuery, Redshift), dbt, orchestration (Airflow, Dagster), and streaming (Kafka, Kinesis). Breadth matters — you'll be picking our stack.
- Strong software engineering fundamentals — you write production code (we're a TypeScript shop, with Python in data/ML workflows), not just orchestration configs.
- Experience mentoring or guiding other engineers — through code reviews, pairing, design feedback, or onboarding.
- Located in San Francisco / Bay Area, or willing to relocate.
Bonus
- Experience leading engineers.
- Experience building or supporting ML/data science workflows (feature pipelines, model inputs/outputs, unstructured data extraction).
- Healthcare standards/technologies: FHIR, HIE, IHE, EHR/EMR, NPI, TEFCA, ADT, HL7, HEDIS, RAF, SNOMED, LOINC, ICD-10, etc.
Benefits
- Competitive equity + compensation package
- Full family Platinum health insurance, dental, and vision coverage
- 401(k) retirement plan + matching
- Flexible work from home or in-office
- Healthy lunches are complimentary when working in-office (and breakfast + dinners as needed)
- Quarterly company off-sites with the team
- MacBook provided by us
- Unlimited PTO (we work hard, but trust you to take time you need to be at your best)
Skills
Spark, Parquet, Iceberg, Delta Lake, Snowflake, BigQuery, Redshift, dbt, Airflow, Dagster, Kafka, Kinesis, TypeScript, Python, AWS
Similar jobs
Data Engineering jobsStaff Software Engineer leading design and development of large-scale batch and real-time data pipelines and ML infrastructure to power GenAI/LLM products and features for Airbnb's Messaging, Notifications, and Connectivity organization. Requires 9+ years experience building production ML systems and cross-functional collaboration.
This staff-level data engineer will architect and operate low-latency market data infrastructure, including feed handling, normalization, distribution, and exchange connectivity. The role requires at least five years of backend engineering experience and strong Java or C++ expertise with high-throughput messaging and market data protocols.
Leads large-scale advertising data ingestion, measurement, and agentic workflow systems, combining deep ad-tech expertise with production LLM experience. Requires 10+ years of engineering experience and technical and people leadership in complex enterprise environments.
Staff Software Engineer responsible for architecting, building, and operating Commure’s data warehouse platform, including CDC, lakehouse, query, transformation, and analytics layers. Requires 6+ years of software engineering experience and broad expertise across modern production data infrastructure.
Senior individual contributor responsible for architecting and scaling production data ingestion systems that integrate complex enterprise sources into reliable datasets. Requires 5+ years of backend engineering experience, strong Python, cloud, Kubernetes, Postgres, and data integration expertise.