Latest Data Engineering jobs
Job results
Builds and scales internal data platform by designing data models, pipelines, and analytics infrastructure to transform raw product/business data into reliable datasets for company-wide decision-making. Partners with stakeholders across Product, Engineering, Finance, Marketing, and Sales.
Build and own scalable streaming, storage, and distributed data-platform systems powering a unified telematics API. The role requires strong Java or JVM-language experience, production platform ownership, and expertise in distributed systems and large-scale batch or streaming data.
Build scalable data infrastructure and analytics pipelines to drive business growth, user personalization, and strategic decisions in sleep fitness technology. Requires 6+ years experience with SQL, Python, DBT, Snowflake, and data architecture expertise.
Builds and maintains data pipelines and privacy-preserving infrastructure for AI economic impact research. Collaborates with researchers and economists using Python, cloud platforms, and LLMs to support scalable economic analysis.
Builds infrastructure and tools for Ramp's Analytics and Machine Learning Platforms, supporting data science lifecycle. Partners with AI and ML engineers; requires Python, workflow orchestrators, cloud platforms, and SQL expertise.
Designs, builds, and optimizes large-scale distributed database systems, pipelines, and performance tooling. The role requires expert database architecture and tuning expertise, AWS and Terraform experience, strong cross-functional leadership, and 14+ years of related experience or an equivalent advanced-degree pathway.
Builds and operates petabyte-scale data platform infrastructure using Kafka, Spark, Flink, and Trino to power real-time ML pipelines and analytics. Requires expertise in distributed systems, stream processing, and systems languages like Rust, Go, or Scala.
Builds and scales reliable data pipelines and infrastructure using AWS and tools like Spark, Kafka, and dbt. Collaborates with teams to deliver data solutions, mentors engineers, and drives data architecture strategy. Requires 6+ years experience.
Build and scale data ingestion infrastructure and EHR integrations to power healthcare products, ensuring reliability for thousands of patients. Requires 5+ years in production systems, 3+ in data platforms, and expertise in Python, AWS, SQL, and big data tools like Spark and Kafka.
Designs and builds scalable ETL/ELT pipelines to harmonize data from ERP systems like SAP, Oracle, and NetSuite into unified models. Requires 5+ years in data engineering, advanced SQL/Python, and ERP data expertise for cross-system integration.
Leads enterprise data architecture strategy, designs multi-cloud integrations, Master Data Management, and AI-enabled data quality frameworks. Requires 7+ years in data architecture/engineering and expertise in modern data platforms.
Senior Data Engineer leads data stack evolution, building scalable pipelines and infrastructure for analytics, reporting, and self-serve tools using Snowflake, dbt, AWS. Requires 7+ years experience, BS/MS, expertise in SQL, Python, orchestration tools.
Lead design and implementation of scalable identity resolution and data governance platforms. Build pipelines for identity data management, ensure privacy compliance, and partner with teams to deliver reliable data services. Requires 5+ years Spark/Scala experience.
Designs and evolves scalable Iceberg-based lakehouse architecture, metadata governance, and security controls for analytics, product, and AI systems. Requires 12+ years experience with Python, SQL, Airflow, and major cloud platforms.
Build and own OnePay’s data infrastructure, streaming transformation pipelines, and reporting systems while partnering with engineering, product, compliance, support, analytics, and data science teams. The role requires 5+ years of experience with Databricks, Spark, Scala, Python, SQL, low-latency pipelines, and AWS.
Develops and maintains software for NCBI's biomedical databases like PubMed and GenBank, implementing bioinformatic algorithms and cloud pipelines for large-scale genetic data. Requires Python proficiency, SQL expertise, Linux scripting, and experience with big data handling.
Staff Data Engineer architects scalable data infrastructure, leads ETL/ELT pipelines and data modeling using Snowflake, dbt, Airflow, and Python. Requires 9+ years experience, mentoring skills, and marketplace background.
Senior Software Engineer on the Product Data Platform team builds and optimizes high-performance data, query, and search infrastructure for scalable backend systems. Requires 7+ years experience with relational databases like Postgres, JVM languages, and distributed systems expertise.
Build and evolve GTM data systems, internal applications, and AI tooling for Sales and Marketing users. The role requires 2+ years in engineering or data, strong Python and SQL skills, cloud experience, and familiarity with generative AI implementation.
Build scalable data pipelines and integrations for an AI-powered financial operating system in healthcare. Requires 5+ years data engineering experience with Python, distributed systems, AI pipelines, and big data technologies; NYC-based hybrid role.
Build and maintain scalable data pipelines and lakehouse infrastructure using PySpark, Databricks, and Airflow on AWS. Partner with Data Science and Engineering teams to enhance data quality, observability, and ML platform support. Requires 4+ years experience with Python, SQL, and cloud data stacks.
Leads architecture and development of large-scale data platforms for AI, including storage, streaming, caching, and indexing. Requires 8+ years experience with databases, streaming tools, Kubernetes, and distributed systems.
Founding Analytics Engineer builds analytics infrastructure from scratch, designs data models and metrics frameworks for healthcare financial data, and creates self-service dashboards. Requires 5+ years experience with SQL, dbt, Python, and cloud data warehouses; NYC-based hybrid role.
Builds and maintains end-to-end data pipelines from raw ingestion to analysis, using Python, SQL, and AI tools to create actionable insights for healthcare clients and internal teams. Requires 4+ years experience handling complex data systems with focus on quality and observability.
Leads technical direction and development of Unity Catalog's governance features for secure data and AI asset management at scale. Requires 15+ years in large-scale distributed systems, deep CS expertise, and strong leadership.
Develops distributed data systems like Apache Spark and Delta Lake at massive scale, ensuring high performance and reliability for exabyte-scale workloads. Requires 8+ years in Java/Scala/C++ and deep distributed systems expertise.
Develop distributed data systems including Apache Spark and Delta Lake to handle big data workloads efficiently. Requires 5+ years in Java/Scala/C++ and expertise in distributed systems.
Senior engineer building distributed data systems like Apache Spark and Delta Lake to handle big data processing, ETL, and data science workloads. Requires 5+ years in Java/Scala/C++ and expertise in distributed systems.
Builds and owns customer data integrations for hospital systems, designing production ETL pipelines and reusable connectors. Requires 4+ years experience with Python, SQL, AWS, and direct customer collaboration in ambiguous environments.
Build full-stack systems, tools, and infrastructure for human feedback collection, AI model alignment, and evaluation. Collaborate with researchers to scale production systems and enhance model safety in a fast-paced environment.
Leads the architecture and performance of high-throughput real-time streaming and Lakehouse data platforms. The role requires 8+ years of software engineering experience, deep Scala or Java and JVM expertise, extensive Kafka experience, and strong AWS capabilities.
Build and maintain scalable data models, pipelines, and Core Data tables to transform raw data into actionable insights. Collaborate with data scientists and business teams using SQL, DBT, Snowflake, Airflow, and Python.
Designs, develops, and maintains large-scale data platforms and pipelines using Java microservices. Leads technical direction, mentors engineers, and ensures data quality, governance, and observability in cloud-native environments. Requires 10+ years experience with 4+ in data pipelines.
Builds analytics infrastructure, frameworks, and tools for monitoring clinical AI/ML product performance, enabling clinical case reviews, and driving cross-functional insights. Requires 5+ years experience with Python, SQL, cloud data platforms, and handling sensitive health data.
Build and optimize scalable data infrastructure and ETL/ELT pipelines supporting a Search API. The role requires 3+ years of data engineering experience, Python, SQL and NoSQL expertise, AWS experience, and a quantitative bachelor's degree.
Owns the semantic and modeling layer of the data warehouse by transforming raw data into trustworthy datasets, reusable metrics, and documented data contracts. The role partners with domain and platform teams to support self-service analytics and efficient warehouse operations.
Architects and builds massive-scale data infrastructure for web crawling, embedding model training, and real-time search, handling hundreds of petabytes. Requires expertise in lakehouse architectures, distributed processing pipelines, and streaming systems like Kafka and Flink.
Designs and owns mission-critical data pipelines to enable decision-making across data science, growth, sales, marketing, and product teams. Requires 5+ years experience with scalable pipelines (preferably Airflow), Python, and advanced SQL.
Build and scale data pipelines to process millions of user votes for AI model evaluation. Partner with researchers to deliver insights via dashboards, ensuring data quality and reliability in a fast-paced environment. Requires 5+ years in data engineering with big data tools.
Builds and scales data pipelines, models, and integrations to provide real-time insights for product, engineering, finance, and GTM teams. Requires 4+ years experience with SQL, Python/JavaScript, dbt/Airflow, and data architecture.
Build and scale data ingestion pipelines and connectors for enterprise SaaS apps, transform unstructured data for AI search and agents, ensure reliability and security at petabyte scale. Requires 3+ years backend/data infrastructure experience with distributed systems.
Builds and leads the development of large-scale distributed data systems and pipelines that power products and organizational decision-making. The role requires 7+ years of data engineering experience, cloud expertise, and strong technical leadership.
Owns full data stack including database architecture, ETL/ELT pipelines, integrations, and product/GTM reporting. Requires 4+ years experience, expert SQL/Python, ETL tools, data modeling, and statistics. Based in SF or NYC.
Designs and scales distributed data infrastructure for large-scale multimodal training and evaluation at OpenAI. Collaborates with researchers to build reliable, high-performance systems handling massive data volumes in a fast-paced environment.
Designs and develops large-scale data applications and pipelines using Java, Spark, Scala, Kafka, Hadoop, and AWS. Requires at least six years of Java development experience focused on data engineering and streaming platforms.
Designs and runs massive-scale data pipelines for ingestion, normalization, enrichment, and delivery across 80M+ companies and 800M+ people. Manages data operations, BPO vendors, partnerships, monitoring, and cost optimization using Python, Dagster, and DuckDB.
Build scalable data infrastructure to ingest and process millions of hardware telemetry data points per second. Requires 7+ years in data engineering, streaming systems like Flink/Kafka, databases like PostgreSQL/Druid, and languages like Go/Rust/Python.
Builds and modernizes data platform infrastructure handling billions of API calls using tools like Kinesis, S3, Athena, and Airflow. Collaborates with customers to incorporate feedback and requires experience with data stacks including Snowflake or Databricks.
Leads development of high-performance data lakes and ETL pipelines for petabyte-scale cyber operations data. Partners with engineers and analysts to build reliable data systems, drives technical initiatives, and mentors team members. Requires 8+ years experience and TS/SCI clearance.
Builds and maintains scalable data pipelines for ingestion, transformation, and reliability using SQL, Python, dbt, and Fivetran to support data science, engineering, and product teams. Requires proven data engineering experience with modern data stack tools.