Skip to content

Data Engineer

Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.

About the job

Responsibilities

  • Build the pipelines that pull every system the company runs on, ERP, ATS, project management, construction software, telemetry, into one queryable layer.
  • Own the data model behind the company's live knowledge graph: entities for sites, equipment, schedules, and people that tools and agents build on.
  • Ship datasets and services with SLAs that internal tools, dashboards, and ML models depend on daily.
  • Turn messy vendor and field data, PDFs, spreadsheets, exports, into structured, trustworthy inputs.

Requirements

  • Built and operated production data pipelines that other teams' products depended on.
  • Modeled a messy real-world domain into schemas that held up as the business changed.
  • Treat data quality as an engineering problem: tests, monitoring, and lineage, not spot checks.
  • Done real work extracting structure from unstructured sources.
  • Move fast with AI tools and modern data stacks without leaving a swamp behind.

Nice-to-Haves

  • Postgres, dbt, or warehouse internals.
  • Streaming and eventing.
  • LLM-based extraction.
  • Construction, manufacturing, or supply chain data.

Skills

Data Pipelines, Data Modeling, Data Quality, ETL, dbt, Postgres, Data Warehousing, Streaming, LLMs, SQL

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Data Operations
$250k+/yrHybridData Engineering

Own end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.

OpenAI

OpenAI

Mountain View, CA
Data Engineer, Monetization Data Platform
$230k+/yrOn-siteData Engineering

Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.

The Voleon Group

The Voleon Group

New York, NY
Software Engineer, Strategy Research Analytics
$230k+/yrRemote3+ YOEData Engineering

Build and evolve reliable analytics infrastructure, pipelines, schemas, and foundational datasets supporting quantitative research across strategies. The role requires strong Python and SQL skills, distributed data-platform experience, and ownership of observability, performance, and reproducibility.

Anthropic

Anthropic

San Francisco, CA
Data Engineer, GTM
$320k+/yrHybrid5+ YOEData Engineering

Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.

Anthropic

Anthropic

San Francisco, CA
Infrastructure Capacity Planner, Demand Planning
$320k+/yrHybridData Engineering

Own medium-range demand forecasts for Anthropic's expanding infrastructure fleet across accelerators, CPU, storage, network, and managed services. The role requires hands-on SQL/Python modeling, large-scale infrastructure planning experience, and partnership with sourcing, Finance, and efficiency teams.