Skip to content
OpenAIOpenAI

Data Engineer, People Innovation Labs

Build and manage data pipelines for people analytics and internal products like OpenHouse at OpenAI's People Innovation Labs. Collaborate with analytics and engineering teams using Databricks, Spark, and ETL tools; requires 3+ years data engineering experience.

About the job

Responsibilities

  • Design, build and manage people data pipelines, ensuring all data is seamlessly integrated into our Databricks warehouse.
  • Develop canonical datasets to track key people metrics and People Innovation Labs product metrics.
  • Work collaboratively with various teams, including Data Platform, Data Science, People Analytics, and Compensation and Equity to understand their data needs and provide solutions.
  • Implement robust and fault-tolerant systems for data ingestion and processing.
  • Participate in data architecture and engineering decisions, bringing your strong experience and knowledge to bear as the primary data engineering expert on the team.
  • Ensure the security, integrity, and compliance of data according to industry and company standards.

Requirements

  • 3+ years of experience as a data engineer and 8+ years of any software engineering experience (including data engineering).
  • Proficiency in at least one programming language commonly used within Data Engineering, such as Python, Scala, or Java.
  • Experience with data warehousing technologies such as Databricks and Snowflake, and expertise with ETL schedulers such as Fivetran, Airflow, Dagster, Prefect, or similar.
  • Experience with distributed processing technologies and frameworks, such as Spark, Hadoop, Flink and distributed storage systems (e.g., HDFS, S3).

Skills

Python, Scala, Java, Databricks, Snowflake, Fivetran, Airflow, Dagster, Prefect, Spark, Hadoop, Flink, Hdfs, S3

Fluidstack

Fluidstack

Austin, TX
Data Engineer
$269k+/yrOn-site5+ YOEData Engineering

Build and own production data pipelines, knowledge graph data models, and structured datasets from messy sources (PDFs, spreadsheets, telemetry) to power internal tools, dashboards, and ML models at a frontier AI compute infrastructure company. Requires experience operating depended-on pipelines, schema modeling, data quality engineering, and unstructured data extraction.

Anthropic

Anthropic

San Francisco, CA
Data Engineer, GTM
$320k+/yrHybrid5+ YOEData Engineering

Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.

Anthropic

Anthropic

San Francisco, CA
Infrastructure Capacity Planner, Demand Planning
$320k+/yrHybridData Engineering

Own medium-range demand forecasts for Anthropic's expanding infrastructure fleet across accelerators, CPU, storage, network, and managed services. The role requires hands-on SQL/Python modeling, large-scale infrastructure planning experience, and partnership with sourcing, Finance, and efficiency teams.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Data Operations
$250k+/yrHybridData Engineering

Own end-to-end data sourcing and vendor operations that help researchers train and evaluate frontier AI models. The role requires strong judgment, communication, problem-solving, and comfort managing ambiguous, fast-changing projects.

OpenAI

OpenAI

Mountain View, CA
Data Engineer, Monetization Data Platform
$230k+/yrOn-siteData Engineering

Build and operate scalable monetization data platforms, pipelines, models, and quality systems spanning product, financial, and operational data. The role partners with Product Engineering, Finance, Accounting, Analytics, and GTM teams to deliver reliable, observable data products.