Latest Data Engineering jobs
Job results
Lead the design, architecture, and implementation of a large-scale facilities telemetry platform for AI data centers, ingesting sensor data from industrial protocols into queryable time-series systems while setting data standards and leading a small team.
Technical leader for Airbnb’s Unified Data Store client stack, responsible for long-term architecture, distributed data access, reliability, and developer experience. Requires 9+ years of industry experience and deep expertise in large-scale distributed systems.
Build and operate high-throughput streaming systems and database internals for Snowflake’s cloud-scale data platform. The role requires 5+ years of distributed-systems or related experience, strong programming skills, and expertise in correctness, fault tolerance, and state management.
Build and evolve Snowflake’s real-time stream processing and data transformation platform, focusing on low-latency execution, correctness, performance, and cloud-scale reliability. The role requires strong distributed systems fundamentals and experience with streaming or transformation engines.
Leads the architecture and delivery of scalable data platforms, pipelines, warehouses, and semantic layers supporting analytics, reporting, product experiences, and AI evaluation. Requires 7+ years in data or analytics engineering and strong expertise in SQL, data modeling, Snowflake, dbt, orchestration, and AWS.
Build and scale secure data infrastructure at Anthropic, including access control systems, financial data pipelines, cloud storage reliability, and analytics tooling. Requires 10+ years building data/distributed systems, 3+ years leading complex projects, and deep experience with cloud/IaC or programming languages.
Build and maintain compliance datasets, data models, ETL pipelines, and measurement frameworks to power BSA/AML detection, monitoring, and controls at Coinbase. Requires 5+ years analytics/data engineering experience, advanced SQL, Python, dbt/Airflow, and regulated environment background.
Build and operate production data pipelines, observability tools, and planning systems to maximize utilization, efficiency, and attribution of Anthropic's large-scale multi-cloud accelerator and CPU fleet. Requires strong Python/SQL, cloud operations, and Kubernetes experience in a high-ambiguity environment.
Forward Deployed Data Engineer building hybrid data pipelines and semantic layers for Hilbert's AI Growth Engine. Implements warehouse-native or managed ClickHouse integrations, partners with AI agents for accelerated onboarding, and ensures reasoning consistency across customer environments.
Senior Platform Engineer designing, building, and scaling database infrastructure for production and AI systems, including relational, analytical, and vector stores. Requires 7+ years experience with distributed systems, high-availability databases, and supporting AI workloads like vector search and RAG.
Senior Data Engineer building scalable data pipelines, services, and AI-ready data layers in Go/Scala/ClickHouse to power internal products, analytics, and agentic AI for go-to-market, engineering, and product teams. Requires 5+ years experience in production data systems, strong programming, SQL, and databases.
Senior Data Engineer responsible for building scalable ingestion pipelines, normalizing, and maintaining large-scale financial and alternative datasets from global vendors to support quantitative research and alpha generation. Requires 5+ years data engineering experience in finance/quant environments, strong Python/SQL/Linux skills, and deep knowledge of market/tick/reference data across asset classes.
Software Engineer building scalable data infrastructure, cataloging, versioning, and lineage tools to support ML research and production workflows at an AI-driven hedge fund. Requires 3+ years experience, strong software design skills, and expertise in a modern language like Python or Java.
Build and operate realtime and batch data pipelines processing billions of events daily at xAI. Design distributed data platforms, own data correctness, create shared datasets for product and business teams, and partner on data acquisition using tools like Spark, Kafka, Flink, and SQL.
Build analytical and BI infrastructure for Scale's Public Sector unit. Develop scalable data pipelines, models, warehouses, and quality tests from ambiguous processes to enable decision-making; requires 5+ years experience, SQL mastery, Python/R, DBT, and active Secret clearance.
Lead the design and implementation of a unified semantic layer and data models from complex enterprise systems (SAP, Salesforce, Workday) to create AI-ready datasets that power intelligent agents, analytics, and decision-making. Requires 10+ years data engineering experience, semantic modeling expertise, and hands-on AI-generated code deployment.
Field Engineer building and deploying data pipelines, integrations, and backend systems for government customers at customer sites. Requires active Secret clearance, Python, ETL, cloud technologies, and >50% onsite availability.
Build scalable financial data pipelines, backend services, and data models supporting Spotify’s accounting and reconciliation systems. The role requires 3+ years of backend or data-intensive application experience with Java or Scala and distributed data processing technologies.
Senior Data Engineer building and scaling Checkr's centralized People Data platform and large-scale batch/streaming pipelines that power identity and people records for all products. Requires 5+ years experience with Python, PySpark, SQL, Spark, Kafka and AWS data services; leads execution, mentors, and partners with cross-functional teams.
Senior Data Engineer building scalable data pipelines, infrastructure, and architecture on AWS using Spark, Metaflow, and orchestration tools. Requires 5+ years data engineering experience with big data technologies; ML/healthcare background is a plus.
Integrations Associate responsible for implementing, maintaining, and optimizing eligibility and claims data feeds for Garner's healthcare platform. Requires analytical skills, familiarity with healthcare file formats (EDI 834/837), SFTP/PGP, and strong cross-functional communication. 0-5 years healthcare data experience.
Build scalable data pipelines, cloud-native infrastructure, and SQL intelligence systems for AI-powered data operations. The role requires 3–5 years of data engineering experience, strong Python and SQL skills, and expertise in query optimization, SQL parsing, and data pipelines.
Build an end-to-end analytics and business intelligence Data Cloud platform at Rippling, replacing customer data lakes, warehouses, and pipelines with integrated ingestion, transformation, lineage, catalogs, and visualization. Develop large-scale data systems using Python, Trino, Iceberg and Temporal; explore ML/LLMs for automated insights.
Build the company’s compute intelligence platform, including warehouse infrastructure, production pipelines, data models, dashboards, and AI-accessible analytics. The role requires 3+ years of data-focused engineering experience, strong Python and SQL skills, and cross-functional business judgment.
Data Engineering Manager to own market data platforms and analytical data systems at a proprietary trading firm. Hands-on role managing a small team while writing production code, leading KDB+/Q time-series architecture, building low-latency pipelines, and evolving the platform for AI-driven consumers. Requires 7+ years data engineering experience with strong KDB+/Q and Python background.
Build and scale the core database infrastructure powering Claude at Anthropic, including data plane/control plane, data movement (CDC, migrations), and caching systems that support millions of users and frontier AI research across multi-cloud environments. Requires deep expertise in distributed databases and production storage systems.
Research Engineers at Distyl build data systems and pipelines that power reliable compound AI workflows in enterprise environments. They create data quality frameworks, synthetic data strategies, and evaluation tools while partnering with researchers and customers to turn raw data into production AI value.
Staff Software Engineer building and scaling Plaid's Data Infrastructure platform (warehouses, lakehouses, Spark, streaming, orchestration). Lead projects to improve ML workflows, data freshness, ETL pipelines; mentor engineers and reduce operational burden. Requires 6+ years software engineering with deep data infrastructure expertise.
Own end-to-end data migrations for government agencies switching to GovWell's AI platform. Clean messy legacy data (SQL/Python), lead customer calls with non-technical staff, ensure high-quality production data, and drive process improvements.
Builds and owns large-scale data processing systems and pipelines for an advertising platform handling billions of requests and over 1 PB of daily data. Requires extensive Java or Scala experience, Big Data expertise, cloud and warehouse proficiency, and a bachelor's degree.
Leads client data conversion projects, transforming historical portfolio data from legacy systems into Addepar while improving migration workflows. Requires at least two years of technology, finance, or consulting experience, Python proficiency, and knowledge of financial products and securities modeling.
Senior Data Engineer on the Risk or Compliance team building data models, pipelines, monitoring, and AI agents to support financial crimes detection, risk decisions, and ML datasets across Cash App, Square, and Afterpay. Requires 8+ years experience, strong SQL/Python/DBT skills, and full lifecycle data engineering expertise.
Senior Backend Engineer building and operating distributed data infrastructure for streaming, batch processing, analytical stores, and warehouse exports at scale. Requires 6+ years backend experience with data pipelines, distributed systems, and production ownership (on-call).
Lead data engineering and analytics engineering teams to design and own ETL/ELT pipelines, data modeling, quality, and governance. Requires 10+ years data engineering experience including 4+ years managing teams, deep expertise in SQL/Python/modern data stack, and partnering with DS/Product/Eng.
Build and own a near real-time data streaming and messaging platform for a logistics company. Design event-driven data pipelines, ingestion, processing, models, and APIs; requires advanced Python, SQL optimization, and real-time data systems experience.
Staff Data Engineer owning end-to-end client data migrations at Machinify. Reverse-engineer legacy healthcare systems, architect and build Spark/Airflow pipelines, ensure data fidelity via automated reconciliation, and coordinate cross-functional teams from discovery through go-live. Requires 8+ years hands-on data engineering, strong Python/SQL/Spark/Airflow, and AI-assisted development experience.
Senior Software Engineer building scalable security data pipelines and streaming APIs at Stripe. Lead design of high-throughput event processing, security signal logic, and developer-friendly platforms; requires distributed systems and data engineering expertise.
Senior Data Engineer owning architecture and technical vision for ClickUp's data platform. Build scalable pipelines on AWS serverless, Snowflake, and dbt; design AI/ML infrastructure; drive cost optimization, standards, and mentor engineers. Requires deep expertise in cloud data systems and large-scale distributed architectures.
Leads day-to-day data center operations, facility relationships, infrastructure deployments, and local technical teams across Western Canada. Requires 5–10+ years of critical-facility experience, MEP knowledge, vendor-management ability, and hands-on leadership.
Build and scale Pigment’s internal finance systems through multidimensional modeling, automated data pipelines, AI agents, and financial workflows. The role requires 3–6 years of FinOps, RevOps, finance analytics, or analytics engineering experience, plus accounting fluency and French and English proficiency.
Leads architecture and hands-on development for Snowflake’s declarative data platform, specializing in streaming ingestion, incremental processing, and dynamic tables. The role requires 14+ years of software engineering experience, deep distributed-systems expertise, and the ability to influence technical direction across organizations.
Software Engineer on the Storage team owning the data layer (databases, caches, scaling strategies) that underpins all Cursor products. Design multi-database architectures, build query guardrails, define storage best practices, and own cache infrastructure for reliability and growth.
Founding Data Engineer to architect Payabli's data platform from scratch: design lakehouse/warehouse, build pipelines, model financial data, and establish governance for a regulated fintech environment.
Design and build scalable big data systems and ETL pipelines using Spark, Kafka, Hive and related technologies. Requires strong data modeling, SQL, and experience with AI coding assistants.
Build and maintain data models, pipelines, and dashboards that power customer experience and compliance operations. Partner with CX and compliance teams to deliver trusted, self-serve analytics.
Builds trusted data models, pipelines, and automated workflows for Linktree’s Strategy and Finance function. The role requires 4–5+ years of data experience, strong SQL and modelling skills, Python automation capability, and effective partnership with non-technical stakeholders.
Build and maintain scalable data pipelines and platforms that enable AI applications to securely access trusted data. Partner with analytics, marketing, and product teams to deliver production-grade data systems.
Principal-level engineer to define and lead Snowflake's core data engineering and streaming primitives (Streams, Tasks, Dynamic Tables) at cloud scale. Requires 15+ years building large-scale distributed data systems and deep expertise in stream processing or data transformation.
Design and maintain scalable data pipelines and lake architecture on GCP/AWS to power analytics, trading tools, and ML initiatives. Requires 5+ years experience, strong SQL/Python, dbt, orchestration tools, and cloud infrastructure experience.
Own the infrastructure connecting massive gameplay data pipelines, GPU clusters, storage, and production inference for action and world models. The role requires substantial hard-infrastructure ownership, hands-on coding, and experience taking systems from design through production.