# Data Lead

**Company:** [Wonderschool](https://hotfix.jobs/companies/wonderschool)
**Location:** San Francisco, CA
**Role:** Data Engineering
**Experience:** 5+ years
**Skills:** SQL, BigQuery, dbt, Data Modeling, entity resolution, deduplication, Data Governance, data quality monitoring, AI Agents, llm data legibility, Data Pipelines, metrics definition
**Posted:** 2026-08-01

> Build and own the canonical data model and pipelines powering Wonderschool's product, government platform, and AI agents. Hands-on role focused on data modeling, identity resolution, state data delivery, governance, and enabling self-serve metrics in a regulated environment.

## Job Description

## Responsibilities
- Build the canonical data model for providers, sites, children, and families across BigQuery, HubSpot, Stripe, and our product databases. One ID, one definition, one source of truth.
- Structure our data so AI agents can read it and act on it: entity tables, snapshots, and pre-computed briefings that power provider coaching at scale.
- Own data delivery for our state government platform: bi-directional Data Hub sync, historical licensing data migration, staging-to-prod loads, validation, and governance sign-off.
- Lead identity resolution and deduplication across legacy state systems so every provider and child has one accurate record.
- Define and document the metrics the business runs on, from "active provider" to churn risk, and make them self-serve for operations and provider success teams.
- Set the data quality bar: monitoring, alerting, lineage, and handling of sensitive data in line with government agreements.
- Use agents to automate your own workflows first, then help other teams do the same with their data questions.

## Requirements
- First or early data hire at a startup.
- Shipped data work in a regulated environment (govtech, healthtech, or fintech), where governance, masking, and audit trails are table stakes.
- Comfortable writing a dbt model as well as working through a data governance agreement with a state counterpart.
- Real scar tissue from entity resolution and dedup work.
- Use AI agents as a daily part of how you work; understand what makes data legible to an LLM.
- Figure things out independently.
- Happiest as a team of one or two and would rather automate a task than hire around it.
- Want to build the system, not inherit one.

## Nice-to-Haves
- Experience with BigQuery, HubSpot, Stripe, product databases.
- Familiarity with dbt, data modeling, identity resolution, deduplication.
- Experience with AI agents (Claude Code, Hermes, OpenClaw) to multiply output.
- Background in government data delivery, state-mandated reporting, or similar regulated data environments.

## Similar roles

- [Senior Analytics Engineer](https://hotfix.jobs/jobs/a3a5e15a-f1e4-4ae5-94c8-1f24fe617a34) - Digible - Remote
- [Senior Manager, Data Engineering](https://hotfix.jobs/jobs/8901dae8-c7af-4e78-9acc-1db5cee94a45) - Farther Finance - New York, NY
- [Senior Analytics Engineer, Product](https://hotfix.jobs/jobs/9c282609-e794-4d6a-9873-b902da9953c9) - Harvey - San Francisco, CA - $156k – $234k/yr
- [Senior Data Engineer II](https://hotfix.jobs/jobs/96b8f8ef-e00f-4f0e-a771-b59989556af3) - Apartment List - Remote - $145k – $207k/yr
- [Senior Analytics Engineer](https://hotfix.jobs/jobs/c5b73bb5-9e7c-4102-a385-841a5637d784) - Alpaca - Remote

**Apply:** https://hotfix.jobs/jobs/78d18ea1-a682-4abe-b120-3eb68becdb1c
**Canonical:** https://hotfix.jobs/jobs/78d18ea1-a682-4abe-b120-3eb68becdb1c