Data Engineer
Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.
About the job
Responsibilities
- Build the pipelines that pull every system the company runs on, ERP, ATS, project management, construction software, telemetry, into one queryable layer.
- Own the data model behind the company's live knowledge graph: entities for sites, equipment, schedules, and people that tools and agents build on.
- Ship datasets and services with SLAs that internal tools, dashboards, and ML models depend on daily.
- Turn messy vendor and field data, PDFs, spreadsheets, exports, into structured, trustworthy inputs.
Requirements
- Built and operated production data pipelines that other teams' products depended on.
- Modeled a messy real-world domain into schemas that held up as the business changed.
- Treat data quality as an engineering problem: tests, monitoring, and lineage, not spot checks.
- Done real work extracting structure from unstructured sources.
- Move fast with AI tools and modern data stacks without leaving a swamp behind.
Nice-to-Haves
- Postgres, dbt, or warehouse internals.
- Streaming and eventing.
- LLM-based extraction.
- Construction, manufacturing, or supply chain data.
Skills
Data Pipelines, Data Modeling, Data Quality, ETL, dbt, Postgres, Data Warehousing, Streaming, LLMs, SQL
Similar jobs
Data Engineering jobsOwn end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.
Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.
Build and evolve reliable analytics infrastructure, pipelines, schemas, and foundational datasets supporting quantitative research across strategies. The role requires strong Python and SQL skills, distributed data-platform experience, and ownership of observability, performance, and reproducibility.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Own medium-range demand forecasts for Anthropic's expanding infrastructure fleet across accelerators, CPU, storage, network, and managed services. The role requires hands-on SQL/Python modeling, large-scale infrastructure planning experience, and partnership with sourcing, Finance, and efficiency teams.