Skip to content
Broccoli AIBroccoli AISan Francisco, CA

Data Analytics Engineer

Build and own the unified data layer and source-of-truth models for a fast-growing AI platform serving home service contractors. Design production pipelines into ClickHouse, create canonical metric definitions, ensure data quality and trustworthiness for customer dashboards, internal analytics, and AI agents.

Salary not listed
On-site4+ YOEData Engineering

About the role

What you'll do

  • Build and run the pipelines. Reliable ingestion from all our sources into ClickHouse — you choose the tooling and own the flow.
  • Model the data. Turn raw feeds into clean, documented tables — including entity resolution, so a customer is the same customer across billing, support, and call data.
  • Build the source-of-truth library. Canonical views and metric definitions that every dashboard and analysis reads from.
  • Make the data Human & AI-ready. Structure our models, definitions, and documentation so both people and AI agents can query them and get the right answer — then build the internal tools that let anyone at Broccoli ask a data question and trust the response.
  • Keep it trustworthy. Freshness checks, quality tests, and alerts — we find out a pipeline broke before a customer does.
  • Run deep dives when the team needs them. Ad-hoc analyses, segment investigations, partner questions.
  • Work closely with engineering. Understand how our systems store and produce data, including schemas, events, and architecture, and give input early on changes so the data that lands in the warehouse is usable, stable, and easy to model.

What we're looking for

  • 4–8+ years in data or analytics engineering — you've built and operated production pipelines end to end, and been the one paged when they broke.
  • Strong SQL and solid Python; hands-on with ETL tooling (Airbyte, Fivetran, Dagster, dbt, or hand-rolled) and orchestration.
  • Real experience with a columnar/OLAP warehouse — ClickHouse ideally; BigQuery, Snowflake, or Redshift transfer fine.
  • Data modeling as a craft: you've designed the tables other people query, and you care what the numbers mean, not just that the pipes run.

Nice to have

  • Self-directed: you've been the first or only data person somewhere, or built a data platform from scratch.
  • ClickHouse specifically — materialized views, performance tuning on event-scale data.
  • Multi-source identity / entity resolution experience.
  • Exposure to customer-facing or multi-tenant analytics (strict customer-level data isolation).
  • B2B SaaS operational data — calls, bookings, jobs, billing — or CRM/field-service data like ServiceTitan.

Skills

SQLPythonETLairbyteFivetranDagsterdbtClickHouseBigQuerySnowflakeRedshiftData Modelingentity resolution
Mercor

Software Engineer, Robotics Data

MercorSan Francisco, CA

Build and own high-scale sensor data pipelines that ingest, process, QC, and deliver multi-modal physical-world data (video, depth, inertial, audio) to train frontier robotics and physical AI models. Requires strong production data engineering experience with video/sensor data at petabyte scale.

130k – 500k/yr
On-site5+ YOEData Engineering
Cerebras Systems

Hardware Analytics Engineer

Cerebras SystemsSunnyvale, CA

Design and optimize hyperscale data pipelines and ML models for multi-terabyte hardware telemetry, performance analysis, anomaly detection, reliability forecasting, and efficiency optimization of AI servers and silicon hardware. Requires Master's degree and 3+ years experience.

214k – 225k/yr
Hybrid3+ YOEData Engineering
Square

Software Engineer, Reconciliation & Reporting

SquareCalifornia

Build and maintain data pipelines for reconciling card-network settlements against internal transaction data to power accurate financial, tax, and regulatory reporting at scale for a global payments platform. Requires 3+ years experience with batch/event-driven pipelines, orchestration tools, and strong SQL/programming skills in a high-stakes financial environment.

121k – 213k/yr
On-site3+ YOEData Engineering
OpenAI

Software Engineer

OpenAISan Francisco, CA

Build and operate reliable, scalable infrastructure and automation for OpenAI's research workloads and data systems (acquisition, processing, ingest, search). Requires strong systems and distributed systems experience, Kubernetes, Linux, networking, and software engineering to improve reliability and reduce operational toil.

255k – 405k/yr
On-site5+ YOEData Engineering
Zoox

Software Engineer, Fleet Simulation

ZooxFoster City, CA

Build and maintain a self-service fleet simulation environment for data scientists and ML engineers to test and evaluate autonomous vehicle orchestration algorithms (dispatch, routing, assignment). Requires production Python experience and building tools for non-specialists.

191k – 266k/yr
On-siteData Engineering