Skip to content
FluidstackFluidstackAustin, TX

Data Engineer

Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.

269k – 317k/yr
On-site5+ YOEData Engineering

About the role

Responsibilities

  • Build the pipelines that pull every system the company runs on, ERP, ATS, project management, construction software, telemetry, into one queryable layer.
  • Own the data model behind the company's live knowledge graph: entities for sites, equipment, schedules, and people that tools and agents build on.
  • Ship datasets and services with SLAs that internal tools, dashboards, and ML models depend on daily.
  • Turn messy vendor and field data, PDFs, spreadsheets, exports, into structured, trustworthy inputs.

Requirements

  • Built and operated production data pipelines that other teams' products depended on.
  • Modeled a messy real-world domain into schemas that held up as the business changed.
  • Treat data quality as an engineering problem: tests, monitoring, and lineage, not spot checks.
  • Done real work extracting structure from unstructured sources.
  • Move fast with AI tools and modern data stacks without leaving a swamp behind.

Nice-to-Haves

  • Postgres, dbt, or warehouse internals.
  • Streaming and eventing.
  • LLM-based extraction.
  • Construction, manufacturing, or supply chain data.

Skills

Data PipelinesData ModelingData QualityETLdbtPostgresData WarehousingstreamingLLMsSQL
Fluidstack

Dev Ops, Facilities Pipeline

FluidstackAustin, TX +3

Build and operate high-frequency facilities telemetry data pipelines (power, cooling, BMS, sensors) for scaling AI data centers. Stand up ingestion/streaming infrastructure, automate deployments with IaC, and own end-to-end reliability for gigawatt-scale operations.

269k – 317k/yr
On-site5+ YOEData Engineering
Fluidstack

Software Engineer, Facilities Pipeline

FluidstackAustin, TX +3

Build and own the facilities data pipeline for AI data center telemetry, including ingestion from industrial protocols (BACnet, Modbus, OPC UA), data quality tooling, and serving clean APIs/datasets for dashboards, controls, and ML. Requires production data pipeline experience with on-call ownership, industrial protocol integration, and full-stack debugging.

269k – 317k/yr
On-site5+ YOEData Engineering
Anthropic

Data Operations Manager, Human Data

AnthropicSan Francisco, CA +1

Data Operations Manager responsible for building and scaling data strategies, vendor partnerships, and high-quality data pipelines to advance frontier AI research in RLHF, safety, tool use, and agentic systems at Anthropic. Requires 3+ years operations/PM experience, strong project management, data analysis skills, and passion for AI safety.

270k – 365k/yr
Hybrid3+ YOEData Engineering
Anthropic

Analytics Data Engineer

AnthropicSan Francisco, CA +2

Builds and manages data pipelines using dbt, SQL, and Python to create scalable analytics infrastructure. Develops dashboards and self-serve tools for company-wide metrics, partnering with Engineering, Product, and GTM teams. Requires 5+ years experience.

275k – 370k/yr
Hybrid5+ YOEData Engineering
OpenAI

Software Engineer, Data Infrastructure - Research

OpenAISan Francisco, CA

Designs and implements dataset infrastructure for OpenAI's large-scale LLM training stack, including standardized APIs for multimodal data, scaling pipelines across GPU fleets, and performance debugging. Requires strong distributed systems experience and collaboration with researchers.

250k – 380k/yr
On-siteData Engineering