Latest Data Engineering jobs
Job results
Designs and implements dataset infrastructure for OpenAI's large-scale LLM training stack, including standardized APIs for multimodal data, scaling pipelines across GPU fleets, and performance debugging. Requires strong distributed systems experience and collaboration with researchers.
Develops and maintains safety data pipelines, dashboards, and AI-assisted triage processes for AV fleet data to support compliance and internal learnings. Requires 4+ years with large-scale datasets, Python/SQL expertise, and dashboard tools like Looker/Databricks.
Builds and scales data pipelines for video AI datasets, owning end-to-end projects from sourcing to delivery with ML filters and dashboards. Requires strong Python skills, clean code practices, and customer communication.
Build and operate a declarative data platform supporting ingestion, transformation, and consumption. The role requires cloud data platform, orchestration, infrastructure-as-code, event-driven architecture, SQL, and distributed systems expertise.
Build and scale analytical data platforms, warehouses, and pipelines supporting customer dashboards and high-volume systems. The role requires 3+ years of backend and infrastructure experience plus expertise in data architecture, schema design, integrations, and scalable data platforms.
Build and maintain core data models, pipelines, and reporting infrastructure using modern data stack tools to enable decision-making across product, ops, and GTM teams. Requires 3-6 years of production data/analytics engineering experience and expert SQL.
Builds and improves large-scale pretraining data pipelines, mixtures, and curation methods for Cohere’s language models. The role combines software engineering and research, requiring Python, data-pipeline development, and experience with large datasets and processing frameworks.
Builds and owns data pipelines, ETL processes, and infrastructure to power reporting and decision-making. Requires 4+ years experience with Python, PostgreSQL, Snowflake, and data modeling.
Leads architecture and evolution of a large-scale data platform spanning ingestion, lakehouse storage, streaming, governance, and access. Requires 10+ years of software engineering experience, deep production expertise with Spark and distributed systems, and strong cloud data-platform experience.
Builds and operates Habitat, OpenAI's core online database platform handling high-QPS, latency-sensitive workloads. Owns end-to-end distributed systems for storage, caching, routing, CDC, and privacy; requires 8+ years experience with Rust/Python expertise.
Builds and maintains scalable data processing pipelines and backend systems for a data curation platform that optimizes training data for ML models. Partners with researchers to integrate research capabilities, ensuring reliability and security for customer data.
Builds scalable data pipelines and infrastructure for AI research, processing petabyte-scale anime data across 10k GPUs. Partners with researchers using distributed systems, big data tools, cloud services, requires 3+ years generalist experience.
Research Engineer owning datasets for training world simulation AI models, designing multimodal datasets, running experiments, and building data pipelines to enhance model capabilities across tasks like robotics and creative tools. Requires 4+ years in ML with experience in generative models and frameworks like PyTorch or JAX.
Builds and operates scalable data infrastructure including compute fleets, storage systems, and streaming platforms to support OpenAI's AI products, research, and analytics. Requires 4+ years in data or infrastructure engineering with expertise in Spark, Kafka, and distributed systems.
Build and manage data pipelines and canonical datasets for product metrics, safety systems, and business decisions. Collaborate with cross-functional teams including Data Science and Research; requires 3+ years data engineering experience with Spark, ETL tools, and distributed systems.
Builds and leads data acquisition systems including web crawling, ingestion, and scalable distributed processing for model training. Requires 4+ years experience, expertise in Kubernetes and large-scale data systems, and BS/MS/PhD in Computer Science.
Builds ETL pipelines and Python scripts to ingest, transform, and manage customer R&D datasets for a scientific web platform. Targets recent graduates with Python scripting, data manipulation experience, and CS coursework.
Build and refine data models using SQL and Python to enable self-service analytics, analyze customer data, and collaborate with cross-functional teams. Requires 5+ years experience in analytics/data roles and strong data modeling skills.