Latest Data Engineering jobs
Job results
Build and maintain production data pipelines that prepare conversational, voice, and multimodal data for ML model training and evaluation. Partner closely with ML engineers to deliver high-quality, versioned datasets and infrastructure.
Senior engineer building and operating Brex’s data platform and infrastructure, partnering with product and analytics teams to deliver data-backed products. Requires 5+ years in data infra/platform roles and experience with Snowflake, Flink, Airflow, dbt, Kafka, and Kotlin/Python.
Lead data strategy and build ETL infrastructure, integrations, and reusable components for internal enterprise systems. Requires 5+ years experience, strong Java/Python, SQL, and API integration skills.
Staff Data Engineer building and scaling data pipelines, integrations, and workflow orchestration systems. Owns architecture, IaC strategy, and technical leadership across large-scale data infrastructure.
Own data and analytics end-to-end: architect internal systems, build metrics/dashboards, and translate customer and product signals into structured inputs for AI agents.
Build and maintain Confido's centralized data warehouse and analytics infrastructure. Design scalable data models, establish data standards, and enable self-service analytics across the organization.
Build and govern scalable Workato integrations, automation frameworks, and data pipelines across enterprise systems. The role requires integration development experience, API expertise, familiarity with enterprise platforms, and exposure to secure AI and LLM implementations.
Architect and build foundational data infrastructure for massive simulation outputs. Design novel data models and high-throughput pipelines to feed LLMs with structured context from complex, state-based environments.
Lead the data engineering team to build and evolve a modern cloud-native data platform, drive data culture, and deliver high-impact data products for risk, customer acquisition, and financial services.
Hands-on Data Engineer building the core data layer for a fast-growing AI observability startup. Own data models, pipelines, and trusted metrics across product usage, revenue, and GTM systems while partnering with Sales, RevOps, Marketing, and Finance.
Build and maintain data infrastructure processing petabytes of data. Own end-to-end projects for data ingestion, transformation, and serving systems. Requires 3+ years of software engineering experience.
Own and extend customer data ingestion platform and large-scale pipelines powering AI workers. Build data lake, retrieval layer, and infrastructure for syncing, enriching, and querying customer data across CRMs and third-party systems.
Senior Data Engineer to own foundational data model, set standards for testing/documentation/alerting, build ingestion pipelines with dbt/Fivetran/Python, and architect governance for a regulated wealth management platform.
Staff-level Data Infrastructure Engineer to architect and evolve the data platform (Snowflake, ingestion, orchestration, CI/CD, AWS infra) serving analytics, product, and ML teams. Requires 10+ years building scalable data platforms and proven technical leadership.
Lead and scale Headway's data engineering team, owning architecture for data warehouse, pipelines, dbt transformations, and orchestration to power analytics, ML, and operations. Requires 8+ years data engineering experience and 3+ years managing teams.
Senior Data Engineer owning end-to-end data domains for industrial plant operations. Designs pipelines, schemas, and contracts from messy sensor/lab sources to support ML and operational decisions.
Lead multi-year vision and architecture for data governance and quality at scale. Define best practices, tooling, and culture while coaching senior engineers and influencing executive strategy on compliance and data integrity.
Lead the technical direction and team for Justworks' core data platform. Own architecture, infrastructure, and standards for pipelines, orchestration, and data governance while managing managers and ICs.
Staff Data Platform Engineer building and leading AWS-native data platform architecture, orchestration, governance, and AI-readiness for analytics and ML workloads. Requires 8-10+ years experience with AWS data systems and strong technical leadership.
Senior engineer responsible for data infrastructure powering product usage tracking, billing, and CRM integrations. Requires 8+ years experience with PHP/Laravel, AWS, relational databases, and React.
Build and operate large-scale multimodal data pipelines for AI avatar model training. Design production-grade systems for petabyte-scale video, audio, and text data.
Staff-level engineer leading design and development of Pinterest’s exabyte-scale data lake storage platform using Iceberg and related big data technologies to support ML/AI workloads.
Senior backend/infrastructure engineer expanding Sentry's time-series data platform (Snuba/ClickHouse) to handle petabyte-scale events with sub-second latency. Requires 4+ years experience and distributed storage expertise.
Senior engineer building data pipelines and tooling to curate large-scale autonomy datasets from drone logs and media for ML model training. Requires 5+ years experience, strong Python/C++ skills, and production data pipeline expertise.
Lead and grow the Trust & Safety Data Engineering team, defining roadmap and technical strategy. Build privacy-safe datasets and pipelines for abuse detection, fraud detection, and safety monitoring. Partner with stakeholders to ensure launch readiness and operational rigor.
Build and operate the identity data platform that ingests, transforms, and serves high-volume identity data to power all Lumos products. Own ingestion pipelines, service layers, APIs, and observability for correctness and reliability.
As a Senior Software Engineer, Data Processing, you will own the data processing layer at ingestion, building and operating systems that transform large-scale source data into clean, structured, AI-ready datasets. This is a hands-on, backend- and data-heavy role with end-to-end ownership of data pipelines.
As a Data Engineering Manager, you will lead the Data & ML Platform team, owning platforms for analytics, experimentation, and machine learning. You will guide the evolution towards a streaming-first, ML-ready architecture and partner with Data Science to operationalize models.
As a Staff Data Engineer, you will architect and scale Imprint's data platform, optimizing infrastructure and driving technical excellence. You will build critical financial reporting pipelines, establish data standards, and mentor other engineers.
As a Member of Technical Staff on the Data Platform team, you will design and operate large-scale batch and streaming data pipelines, lead data orchestration architecture, and build self-serve data platforms. This role involves shaping the technical direction of Perplexity’s data ecosystem and mentoring engineers.
Quantitative Developer building and maintaining portfolio compliance and analytics features. Works cross-functionally to onboard clients, debug calculations, and implement new functionality. Requires strong Python/Java/C# skills and quantitative background.
As a Software Engineer, Data Infrastructure, you will design and build resilient and scalable data systems for real-time data movement, processing, and storage. You will drive innovation in data ingestion, transformation, and delivery, enabling actionable insights across the company.
As a Software Engineer on the Data & Analytics Platform team, you will design, build, and optimize the data platform to support various data-driven initiatives. You will work with cross-functional teams to architect scalable solutions and implement data infrastructure using modern data technologies.
As a Senior Staff Data Platform Engineer, you will define platform standards, lead cross-domain initiatives, and shape the future of Zocdoc's analytics and data infrastructure. You will ensure the data ecosystem is secure, reliable, compliant, performant, and cost-efficient.
Build and operate large-scale web data pipelines for pretraining language models, including extraction, filtering, quality scoring, deduplication, and corpus analysis. The role requires strong Python and data-engineering skills, experience with large web datasets, and collaboration across research and engineering teams.
As an IT Controls Data Engineer, you will build and maintain data infrastructure for audit readiness, IT controls, and continuous control monitoring. This role involves designing pipelines, datasets, and automated validation to ensure reliable control data.
As Head of Data, you will lead the data function end-to-end, shaping strategy, building data pipelines, and driving product and business decisions. You will also build and lead a small, high-leverage team, pushing the boundaries of AI tooling.
As a Data Engineer, you will build and maintain data pipelines, dbt models, and infrastructure on AWS and Snowflake. You will partner with BI/Analytics Engineering, take operational responsibility, and mentor junior team members.
Senior data engineer building scalable data products, models, and governance processes. Focus on automation, performance optimization, and enabling self-serve analytics for a fast-growing SaaS company.
Build and own high-performance JVM-based connectors, drivers, and integrations connecting ClickHouse with streaming frameworks and the broader data ecosystem. The role requires 6+ years of software development experience, deep Java expertise, and production experience with streaming connectors and distributed messaging systems.
Build and maintain high-performance Java/JVM connectors, drivers, and SDKs that integrate ClickHouse with streaming and data-processing ecosystems. The role requires 6+ years of software development experience, production connector expertise, and strong knowledge of Java concurrency, distributed messaging, and database systems.
Build and maintain high-performance Java/JVM connectors, drivers, and SDKs that integrate ClickHouse with streaming and data-processing ecosystems. The role requires 6+ years of software development experience, strong Java concurrency and performance expertise, and production experience with streaming connectors.
Senior software engineer building and maintaining high-performance JVM-based connectors, drivers, SDKs, and integrations for ClickHouse’s streaming and data engineering ecosystem. Requires 6+ years of experience with Java, distributed messaging, streaming frameworks, concurrency, and scalable data integration systems.
Build and maintain high-performance Java/JVM connectors, drivers, SDKs, and streaming integrations for ClickHouse’s data ecosystem. The role requires 6+ years of software development experience plus deep expertise in Java concurrency, distributed messaging, streaming frameworks, and data-intensive systems.
Senior software engineer building and maintaining high-performance Java/JVM connectors, drivers, SDKs, and integrations for ClickHouse across streaming and data-processing ecosystems. Requires 6+ years of software development experience and deep expertise in Java concurrency, distributed messaging, databases, and performance optimization.
Own the technical vision and architecture for a data & ML platform serving analytics, product, and machine learning workloads. Drive cross-org initiatives, set platform standards, and build infrastructure at the intersection of data engineering and ML systems.
Build and scale reliable data pipelines for the Growth team, establish dbt and data-quality standards, and enable self-service analytics across cross-functional teams. The role requires strong data engineering experience, expert dbt knowledge, and proficiency in Python and SQL.
Lead data infrastructure for a national security tech company building a high-performance data lake and ETL pipelines for petabyte-scale cyber operations datasets. Requires 8+ years experience, strong data lake expertise, and proven technical leadership.
Architect and build Voleon’s batch and realtime streaming platform that powers ML research and production trading systems. Requires 10+ years experience building scalable data infrastructure with Python, Go, and distributed systems.
Builds and maintains petabyte-scale data storage infrastructure for AI training workloads. Requires 4+ years in data infrastructure, Python, Kubernetes, and distributed processing frameworks like Spark or Beam.