Skip to content
AnthropicAnthropic

Engineering Manager, Research Data Platform

Leads technical direction for a research data platform, building scalable pipelines, platform components, and canonical datasets used by ML researchers. The role requires experience with data-intensive systems, schema design, cross-team technical leadership, and hands-on software development.

About the job

Responsibilities

  • Work directly with researchers and supporting engineers to understand workflows, identify high-leverage opportunities, and shape the team’s roadmap.
  • Set technical direction across research data platform components and datasets.
  • Design and build platform components, including libraries, services, and interfaces such as metrics libraries integrated with training frameworks.
  • Own core datasets end to end, including pipelines, schemas, documentation, and data-quality guarantees.
  • Drive adoption of canonical datasets, including the core data model for reinforcement-learning transcripts.
  • Lead complex, multi-quarter projects spanning multiple systems and teams while remaining hands-on in the code.
  • Raise the team’s technical bar through design reviews, mentorship, and high-quality implementation.

Requirements

  • Experience building and operating data-intensive systems at scale, including pipelines, storage layers, and query systems.
  • Strong data-modeling and schema-design skills.
  • Experience setting technical direction or owning the architecture of a data platform used by other teams.
  • Ability to conduct user discovery, iterate with internal users, and measure success through adoption.
  • Ability to build stable interfaces and trustworthy data while research use cases evolve.
  • Ability to lead through influence and align engineers and stakeholders without relying on formal authority.
  • Pragmatic, results-oriented approach and willingness to handle high-leverage operational work.
  • Interest in learning the fundamentals of machine-learning research; deep ML expertise is not required.
  • Interest in the societal impacts of the work.
  • Bachelor’s degree in a relevant field, or an equivalent combination of education, training, and experience.

Nice-to-haves

  • Experience with large-scale ETL and columnar or analytical storage, such as Spark, BigQuery, ClickHouse, DuckDB, or Parquet.
  • Experience with metrics, experiment-tracking, or high-volume time-series systems.
  • Experience with dataset management, cataloging, or lineage tooling.
  • Experience building developer tooling or internal data platforms for demanding technical users.
  • Experience in quantitative trading or similarly exploratory data environments.
  • Working knowledge of machine learning.
  • Experience working in or closely with an ML research lab.
  • Interest in or experience with people management and growing engineers.

Compensation

  • Annual salary range: $405,000–$850,000 USD.
  • Hybrid policy: Staff are currently expected to work from an office at least 25% of the time; some roles may require more office time.

Skills

Data Modeling, Schema Design, ETL, Spark, BigQuery, ClickHouse, Duckdb, Parquet, Metrics Systems, Experiment Tracking, Time-Series Data, Data Cataloging, Data Lineage, Machine Learning, Python

Anthropic

Anthropic

San Francisco, CA
Data Engineer, GTM
$320k+/yrHybrid5+ YOEData Engineering

Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.

Anthropic

Anthropic

San Francisco, CA
Infrastructure Capacity Planner, Demand Planning
$320k+/yrHybridData Engineering

Own medium-range demand forecasts for Anthropic's expanding infrastructure fleet across accelerators, CPU, storage, network, and managed services. The role requires hands-on SQL/Python modeling, large-scale infrastructure planning experience, and partnership with sourcing, Finance, and efficiency teams.

Fluidstack

Fluidstack

Austin, TX
Data Engineer
$269k+/yrOn-site5+ YOEData Engineering

Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Data Operations
$250k+/yrHybridData Engineering

Own end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.

OpenAI

OpenAI

Mountain View, CA
Data Engineer, Monetization Data Platform
$230k+/yrOn-siteData Engineering

Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.