Latest Data Engineering jobs
Job results
Build and manage data pipelines and canonical datasets for experimentation platform, tracking product metrics like user growth and revenue. Collaborate with cross-functional teams at OpenAI; requires 3+ years data engineering experience with Spark, ETL tools, and distributed systems.
Builds and maintains reliable data pipelines and infrastructure using SQL, Python, Airflow, and Redshift to support analytics for Risk, Marketing, Finance stakeholders. Leads projects, reviews code, and scales platform as business grows.
Build and manage data pipelines for people analytics and internal products like OpenHouse at OpenAI's People Innovation Labs. Collaborate with analytics and engineering teams using Databricks, Spark, and ETL tools; requires 3+ years data engineering experience.
Builds and manages data pipelines using dbt, SQL, and Python to create scalable analytics infrastructure. Develops dashboards and self-serve tools for company-wide metrics, partnering with Engineering, Product, and GTM teams. Requires 5+ years experience.
Designs and leads production data pipelines integrating EHR clinical data from Epic Cerner Athenahealth via FHIR HL7 CCDA enforcing healthcare standards for low-latency insights. Builds AWS cloud infrastructure ETL workflows with Python Docker Kubernetes for ML enablement and HIPAA compliance requiring 5+ years healthcare data engineering experience.
Leads Scientific Data Engineering team to build data pipelines, schemas, and parsers for pre-clinical lab data using Python/SQL/AI. Architects solutions, mentors juniors, and delivers customer-focused data products with dashboards in React/Streamlit. Requires 8+ years experience.
Builds scalable data infrastructure for ML training and evaluation in autonomous driving, including pipelines for logs, annotation tools, dashboards, and monitoring systems. Requires Python proficiency, 1+ years experience with large-scale data systems, and CS/EE bachelor's degree.
Builds scalable data infrastructure for autonomous driving ML systems, processing large-scale batch/streaming data from logs and simulations. Requires Python/C++ proficiency, data pipeline experience, and bachelor's in CS/EE; 1+ years experience.
Build scalable data infrastructure for ML training and evaluation in autonomous driving, including batch/streaming pipelines, storage systems, dashboards, monitoring, data mining, and annotation tools. Requires 4+ years experience, Python proficiency, and engineering leadership.
Build scalable data platforms for autonomous driving ML systems, including batch/streaming pipelines, storage, dashboards, and monitoring. Requires 4+ years experience in large-scale data systems, Python/C++, and engineering leadership.
Designs and evolves foundational payroll data models and materialization pipelines across multiple countries. The role requires 5+ years of software engineering experience, strong backend and data modeling expertise, and the ability to lead cross-team architecture and reliability efforts.
Owns end-to-end analytics engineering including dbt models, orchestration, metrics, documentation, and data quality in Snowflake. Requires 5+ years experience, strong SQL/data modeling, and orchestration tooling expertise.
Lead RWD strategy and data engineering at a tech-driven pharma company. Own sourcing, harmonization (OMOP), infrastructure, and vendor management to deliver analysis-ready datasets supporting drug development and portfolio decisions.
Leads design and evolution of analytics infrastructure for research reporting, owning pipelines, datasets, and platform standardization. Collaborates with data scientists using Python, SQL, distributed engines like Presto/Spark. Requires 6+ years in data infrastructure.
Builds and maintains a high-quality data platform delivering Forge data to clients, implementing features with agile methodologies. Requires 3-5+ years in TypeScript/C#/Java/Python, React, modern data architectures, CI/CD, AWS, and SQL databases; hybrid in Soho, NY.
Builds and maintains a high-quality data platform for delivering Forge's financial data to clients, implementing features with agile methods. Requires 3-5+ years in TypeScript/C#/Java/Python, React, modern data architectures, CI/CD, AWS, and SQL/PostgreSQL.
Builds automation, tools, and frameworks for data governance, quality, and compliance using Python and Terraform. Partners with engineering teams to implement trust signals, scorecards, and AI-enhanced workflows for reliable data at scale. Requires 5+ years experience.
Senior Data Engineer building scalable pipelines to transform EHR and claims data into analytics-ready assets while conducting hands-on real-world evidence analyses. Requires 5+ years experience with 2+ years in healthcare data, strong SQL/Python, Snowflake/dbt/Dagster, and familiarity with OMOP and causal inference frameworks.
Builds and scales data pipelines and infrastructure powering real-time AI agents, handling ingestion from CRM, transcripts, and signals into clean representations. Requires 5+ years in data systems, proficiency in Python/SQL/DBT/warehouse tech, and real-time streaming experience.
Build and scale Blink’s data platform by developing reliable ingestion pipelines, workflow orchestration, observability, and analytics infrastructure. The role requires strong SQL and Python, modern data-stack experience, cloud tooling familiarity, and the ability to collaborate across engineering and business teams.
Designs and leads unified data architecture integrating vendor datasets for quantitative research, simulation, and alpha generation across asset classes. Requires 7+ years experience in data engineering, Python proficiency, and financial data modeling expertise.
Builds and scales ETL pipelines, designs data schemas, and owns data quality/governance for 10x growth. Requires 5+ years in data pipelines with SQL, Spark, Airflow, Python, and MPP databases like Snowflake/Redshift.
Builds and scales data infrastructure across the full lifecycle (collection, ingestion, storage, querying) using open-source technologies like Spark and Kafka. Supports cloud, hybrid, and on-prem deployments while collaborating across business units. Requires 3+ years experience and Bachelor's degree.
Builds large-scale data processing pipelines and ML infrastructure to automate data curation, model training, and iteration for autonomous vehicles using real-world and simulation data. Requires 3-5 years experience in data/ML infra, Python, and frameworks like Spark/Airflow/Kafka.
Builds and optimizes big data pipelines processing billions of signals for technographic data services. Requires 5+ years experience with Spark, Hadoop, Java/Scala, and distributed systems; collaborates with data scientists on ML integration.
Builds and scales petabyte-scale data ingestion pipelines for observability platform using Go/C++ on AWS/Azure. Requires 5+ years in distributed systems, strong systems programming, and cloud experience.
Builds scalable financial data infrastructure and AI-powered automation for finance operations, integrating tools like dbt, Snowflake, and agentic workflows to replace manual processes. Requires 3+ years in finance ops/data engineering, SQL/Python proficiency, and finance domain expertise.
Build and maintain scalable data pipelines, analyze large datasets, and develop models/tools to drive product decisions for a code review platform. Requires 4+ years data engineering experience, strong SQL, cloud data services, and visualization tools.
Staff engineer leading Data Platform initiatives like scaling, stream processing, and GRC at massive scale. Requires 7+ years experience in distributed systems and data infrastructure, with strong architectural leadership and cross-functional influence.
Builds scalable data ingestion pipelines, processing systems, and tooling to support AI/ML research and trading operations at a quantitative finance firm. Requires 5+ years experience in robust software engineering with modern languages and data infrastructure expertise.
Staff Data Engineer architects and delivers scalable data products from healthcare datasets, designs high-performance processing systems using SQL, Spark, Python, and AI workflows, and leads cross-functional initiatives for reliable data serving to customers and applications.
Designs and owns canonical data foundations, ingestion pipelines, and AI-ready schemas for financial AI systems in wealth management. Requires 5+ years data platform experience, SQL/Python expertise, custodial data knowledge, and AWS proficiency.
Builds semantic layers with dbt/SQL, agent tooling for autonomous data querying across Athena/Salesforce, and infrastructure for AI-driven insights on usage-based SaaS metrics. Requires 5+ years in data/AI engineering, SQL/Python mastery, and RAG/agentic systems experience.
Senior Software Engineer on the Data Platform team owns data infrastructure, builds scalable ETL/ELT pipelines, optimizes storage, and collaborates with product and data science teams. Requires 5+ years experience in data engineering with Python/Golang, Spark, AWS, and Kubernetes.
Staff Data Engineer owns and evolves data platforms including warehouse architecture, pipelines, and modeling to enable scalable analytics and self-service insights. Requires 7+ years experience, advanced SQL/Python, and expertise with managed data warehouses like Snowflake.
Build production systems and data pipelines that turn evaluation signals into durable datasets, replayable product simulations, and trusted verdicts for Search and Product teams. The role requires 3+ years of software engineering experience, Python and SQL proficiency, distributed data systems expertise, and AWS or lakehouse experience.
Builds measurement, evaluation, and feedback infrastructure for AI agent quality. Designs datasets, pipelines, dashboards, and analysis tools to improve agent reliability, partnering with research and product teams. Requires strong data acumen and software engineering skills.
Leads architecture and strategy for data warehouse, analytics tools, and governance at massive scale. Requires 12+ years experience with data platforms like Spark/Trino and AI tools for productivity.
Builds and maintains enterprise DataOps platform using Kubernetes, cloud services, and data processing tools. Requires 7+ years experience, strong coding skills, and expertise in containerization, GCP/AWS/Azure, Kafka, and Airflow.
Builds and optimizes big data pipelines processing billions of signals for technographic data services, designs scalable architectures, and collaborates with data scientists on ML integration. Requires 8+ years experience with Hadoop, Spark, Java/Scala, and large-scale distributed systems.
Builds and scales distributed data infrastructure powering analytics, AI/ML, and BI at Figma. Requires 5+ years backend experience with batch/streaming tech like Spark, Kafka, and expertise in Snowflake, Golang/Python.
Builds and owns C++-based perception data infrastructure for autonomous vehicles, handling massive 3D sensor data (Lidar, Camera, Radar), validation, and ML pipelines. Requires 4+ years experience, C++ fluency, Python, and systems engineering in ambiguous robotics environments.
Builds and maintains scalable data models and pipelines using Snowflake, dbt, and Databricks to support analytics for patient intake funnel. Leads BI solutions, self-service tools, and best practices with 5+ years experience, advanced SQL, and cross-functional collaboration.
Builds and scales data pipelines for video generation models, including ingestion, annotation via MTurk/Prolific, preprocessing, and curation using Python, AWS, Kubernetes. Requires 3+ years in ML/data engineering, PyTorch experience, and cross-functional collaboration.
Leads design and implementation of scalable security data pipelines and API-first services to deliver real-time security telemetry for enterprise customers. Requires 8+ years in data engineering, distributed systems, and languages like Java, C++, or Rust.
Owns end-to-end data strategy including sourcing, curating, and structuring multimodal data (text, video, images) for AI model training. Requires strong Python, SQL, large-scale processing, and ML-first mindset with LLM experience.
Designs and builds scalable batch/streaming data pipelines for identity verification products, owning end-to-end data initiatives using cloud-native tech. Requires 5+ years data engineering with Spark, AWS, Python/SQL; streaming/orchestration experience preferred.
Leads architecture and implementation of healthcare payer data integrations (EDI X12, HL7/FHIR) and scalable pipelines for AI/ML systems. Requires 5+ years data engineering experience with payer data and team leadership.
Build distributed data storage and processing systems for big data workloads including Apache Spark, Delta Lake, and performance optimization. Requires 5+ years in Java/Scala/C++, strong algorithms knowledge, and distributed systems experience.
Designs and maintains near real-time data streaming platforms, data pipelines, and event-driven architectures. Requires advanced Python, SQL expertise, and 3+ years in software engineering focused on data systems.