Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.
269k – 317k/yr
On-site5+ YOEData Engineering
About the role
Responsibilities
Build the pipelines that pull every system the company runs on, ERP, ATS, project management, construction software, telemetry, into one queryable layer.
Own the data model behind the company's live knowledge graph: entities for sites, equipment, schedules, and people that tools and agents build on.
Ship datasets and services with SLAs that internal tools, dashboards, and ML models depend on daily.
Turn messy vendor and field data, PDFs, spreadsheets, exports, into structured, trustworthy inputs.
Requirements
Built and operated production data pipelines that other teams' products depended on.
Modeled a messy real-world domain into schemas that held up as the business changed.
Treat data quality as an engineering problem: tests, monitoring, and lineage, not spot checks.
Done real work extracting structure from unstructured sources.
Move fast with AI tools and modern data stacks without leaving a swamp behind.
Nice-to-Haves
Postgres, dbt, or warehouse internals.
Streaming and eventing.
LLM-based extraction.
Construction, manufacturing, or supply chain data.
Skills
Data PipelinesData ModelingData QualityETLdbtPostgresData WarehousingstreamingLLMsSQL
Build and operate high-frequency facilities telemetry data pipelines (power, cooling, BMS, sensors) for scaling AI data centers. Stand up ingestion/streaming infrastructure, automate deployments with IaC, and own end-to-end reliability for gigawatt-scale operations.
269k – 317k/yr
On-site5+ YOEData Engineering
Software Engineer, Facilities Pipeline
FluidstackAustin, TX +3
Build and own the facilities data pipeline for AI data center telemetry, including ingestion from industrial protocols (BACnet, Modbus, OPC UA), data quality tooling, and serving clean APIs/datasets for dashboards, controls, and ML. Requires production data pipeline experience with on-call ownership, industrial protocol integration, and full-stack debugging.
269k – 317k/yr
On-site5+ YOEData Engineering
Data Operations Manager, Human Data
AnthropicSan Francisco, CA +1
Data Operations Manager responsible for building and scaling data strategies, vendor partnerships, and high-quality data pipelines to advance frontier AI research in RLHF, safety, tool use, and agentic systems at Anthropic. Requires 3+ years operations/PM experience, strong project management, data analysis skills, and passion for AI safety.
270k – 365k/yr
Hybrid3+ YOEData Engineering
Analytics Data Engineer
AnthropicSan Francisco, CA +2
Builds and manages data pipelines using dbt, SQL, and Python to create scalable analytics infrastructure. Develops dashboards and self-serve tools for company-wide metrics, partnering with Engineering, Product, and GTM teams. Requires 5+ years experience.
275k – 370k/yr
Hybrid5+ YOEData Engineering
Software Engineer, Data Infrastructure - Research
OpenAISan Francisco, CA
Designs and implements dataset infrastructure for OpenAI's large-scale LLM training stack, including standardized APIs for multimodal data, scaling pipelines across GPU fleets, and performance debugging. Requires strong distributed systems experience and collaboration with researchers.