Latest Data Engineering jobs
Job results
Leads the data engineering function and owns the platform supporting custodial ingestion, security master data, canonical financial models, governance, and AI-ready products. Requires 8+ years in data platforms, management experience, broker/dealer exposure, and strong PostgreSQL, AWS, and Python expertise.
The Senior Data Engineer will build and operate trusted pipelines, governed metrics, and self-service analytics for clinical and operational data. The role requires 5+ years of production data engineering experience, strong Snowflake, SQL, and Python skills, and comfort owning data quality end to end.
Leads the design, development, and reliability of foundational data storage, streaming, caching, indexing, and warehousing platforms. The role requires at least five years of backend engineering experience, distributed-systems expertise, and fluency with modern data and cloud technologies.
The Senior Analytics Engineer will build product event data models, pipelines, and semantic layers that enable reliable self-serve analytics. The role requires 5+ years of relevant experience, strong SQL, dbt, Python, and Snowflake expertise, and close collaboration with Product, GTM, Finance, and executive stakeholders.
Build and scale Oracle Fusion financial systems across accounting, tax, treasury, procurement, and FP&A. The role combines financial data architecture, SQL/PL/SQL integrations, workflow automation, and AI-assisted development in partnership with business and engineering teams.
Build and govern the cloud data infrastructure that powers AI skill mining, including BigQuery, storage, IAM, Vertex AI pipelines, and automated data workflows. The role partners with AI, data science, and product teams to deliver observable, privacy-conscious infrastructure.
Owns the analytics transformation layer by building scalable dbt and SQL data models, improving warehouse performance, and establishing data quality standards. Requires 4+ years in analytics or data engineering, strong Python and semantic-layer experience, and proficiency with cloud and data tooling.
Build internal applications, automated workflows, and ETL pipelines that support Commure’s global operations. The role requires strong SQL, JavaScript, Python, full-stack development, database architecture, and CI/CD experience, with at least three years in software, analytics engineering, or technical operations.
Build and own the canonical data model and pipelines powering Wonderschool's product, government platform, and AI agents. Hands-on role focused on data modeling, identity resolution, state data delivery, governance, and enabling self-serve metrics in a regulated environment.
Build and operate highly available datastore infrastructure and platform tooling for Auth0’s distributed systems. The role requires 5–8 years of software development experience, infrastructure and cloud expertise, and familiarity with databases, reliability, and operations.
Staff Analytics Engineer owning centralized data assets, dbt/Snowflake modeling, and governance at a high-growth YC-backed SaaS company. Requires 8+ years analytics engineering experience, deep expertise in Snowflake, dbt, advanced SQL, and SaaS subscription metrics to deliver trustworthy data models that drive business decisions.
Build and own Virta Health's data engagement platform that powers Member Marketing and Coverage Eligibility decisions. Design observable, correct data contracts and modernize legacy integrations into event-driven systems while leading technical delivery in a regulated healthcare environment.
Senior Data Engineer building and scaling clinical data pipelines, lakehouses, warehouses, and ETL/ELT systems for millions of patients to power real-time healthcare products and AI/ML workflows. Requires 6+ years data engineering experience, strong software engineering fundamentals, and mentoring skills.
Lead the design, architecture, and scaling of Metriport's data platform for ingesting and processing real-time clinical data from millions of patients. Own end-to-end data projects, mentor engineers, support AI/ML workflows, and eventually lead a team while staying hands-on. Requires 8+ years building large-scale data platforms with modern cloud-native tools.
Senior Software Engineer building large-scale distributed data systems and pipelines on Scala, Spark, and Databricks to power blockchain analytics and intelligence products that fight financial crime. Requires hands-on Scala/functional programming experience and expertise designing scalable batch/streaming data platforms in the cloud.
Staff Software Engineer on the Data Platform team defining technical direction for large-scale data infrastructure powering fraud detection. Own design of batch/streaming pipelines, set engineering standards, mentor juniors, and partner cross-functionally on scalable, reliable systems in AWS.
Build and own blockchain network infrastructure, SDKs, and platform APIs that abstract complexity for Coinbase's data platform users. Lead initiatives on chain integrations, data migrations, and distributed systems while maintaining APIs, SLOs, and on-call responsibilities. Requires 5+ years software engineering experience with distributed systems and blockchain/crypto infrastructure.
Build, scale, and optimize data pipelines and infrastructure that transform raw events into actionable intelligence for analytics, BI, and business decisions. First dedicated data engineer on a small data team; own ingestion, transformation, semantic layer, and mentor analysts/scientists. Requires 5+ years software engineering with data emphasis, expert SQL, and data warehouse experience.
Lead the design and scaling of LeafLink's data platform and analytics infrastructure. Build reliable pipelines and data models using Airflow, dbt, Python, and AWS to support decision-making across the organization. Requires 8+ years building production data pipelines.
Lead and grow a team of data engineers responsible for Dropbox's core data platform pipelines, self-serve analytics, data quality, reliability, and cost efficiency. Requires 8+ years data engineering experience and 3+ years managing teams, with deep expertise in modern data stacks.
Lead the design, architecture, and development of scalable data infrastructure and platforms at Twilio. Requires 10+ years software engineering experience including technical leadership, deep expertise in big data technologies (ClickHouse, Kafka, Spark), cloud (AWS), and infrastructure-as-code tools.
Build and own data foundations powering model training, product development, and analytics at Parallel. Design scalable ingestion pipelines, storage/serving layers, quality/lineage/observability systems, and anticipate scaling needs in a fast-growing AI infrastructure company.
The Senior Database Administrator will administer AWS-hosted PostgreSQL databases, lead query and schema optimization, guide data architecture, and mentor engineers. The role requires 7+ years managing enterprise databases, strong SQL expertise, and experience with cloud database services and AI-assisted technical workflows.
Staff Marketing Technology Engineer who owns a BigQuery/dbt data platform and builds secure AI applications and agents on Google Cloud. Requires 5+ years of data engineering experience plus strong Python, SQL, cloud deployment, software engineering, and data governance expertise.
Develops and optimizes petabyte-scale cloud database systems, focusing on high-performance data processing, query optimization, and scalability solutions. Requires 2+ years experience, fluency in Java or C++, strong CS fundamentals, and onsite work in Menlo Park or Bellevue.
Lead the unified Data & AI engineering function at BuildOps. Own data platforms, pipelines, ML infrastructure, governance, and a high-performing team to power AI products and establish BuildOps as the trusted system of record for commercial contractors. Requires 10+ years experience scaling data teams and deep expertise in modern data/ML platforms.
Senior Data Platform Engineer owning data infrastructure for identity and fraud detection products. Build scalable ETL/ELT pipelines, data observability, and storage layers using Python/Golang, Spark, AWS, and databases. 5+ years experience required; mentor juniors and collaborate with product/DS teams.
Build and scale Addepar's Analytics Platform using the Data Lakehouse to optimize data analytics, workflows, Ops, and Data Governance. Requires 5+ years experience in platform development for data engineering outcomes, strong Java/Python skills, cloud platforms, CI/CD pipelines, IaC, and observability tools.
Build and operate scalable backend systems, data pipelines, storage solutions, and architectures supporting healthcare advertising products, analytics, and machine learning workflows. The role requires at least one year of data engineering experience, cloud proficiency, and strong Java, SQL, and Python skills.
Build and maintain data pipelines for reconciling card-network settlements against internal transaction data to power accurate financial, tax, and regulatory reporting at scale for a global payments platform. Requires 3+ years experience with batch/event-driven pipelines, orchestration tools, and strong SQL/programming skills in a high-stakes financial environment.
Build and operate reliable, scalable infrastructure and automation for OpenAI's research workloads and data systems (acquisition, processing, ingest, search). Requires strong systems and distributed systems experience, Kubernetes, Linux, networking, and software engineering to improve reliability and reduce operational toil.
Lead design and optimization of scalable data platforms, real-time streaming pipelines, and cross-platform data integrations for supply chain visibility at enterprise scale. Requires advanced Python expertise, technical leadership, and 5+ years software development experience (Staff-level).
Build and maintain a self-service fleet simulation environment for data scientists and ML engineers to test and evaluate autonomous vehicle orchestration algorithms (dispatch, routing, assignment). Requires production Python experience and building tools for non-specialists.
Senior Performance Engineer improving how data products interact with large, complex healthcare datasets. Uses data science skills to optimize performance across data architecture, queries, APIs, and UX for faster, scalable products.
Lead and hands-on architect Current's Analytics Engineering function as a player-coach. Own data warehouse strategy, dimensional modeling in dbt/BigQuery, BI with Looker, team management, and stakeholder alignment in a fintech environment. Requires 10+ years experience including 5+ in leadership.
Build and operate large-scale backend data pipelines and ETL/ELT workflows for acquiring, validating, and storing high-volume business data using Java, Spark, Airflow, Beam, Kafka and GCP services. Requires 3+ years experience in production data engineering.
Build and maintain scalable data pipelines, warehousing, dashboards, and analytical tooling to support Anthropic's Safeguards team in monitoring AI models, detecting abuse, and ensuring safety at scale. Requires strong SQL/Python, modern data stack experience, and collaboration with engineers, data scientists, and policy teams.
Data Platform Engineer responsible for architecting and managing production data pipelines in Snowflake and Databricks, building ETL processes, scaling Terraform deployments, and advancing data governance. Requires 3+ years data engineering experience, strong API and pipeline skills, and comfort in ambiguous startup environments.
Staff Data Engineer building and evolving Checkr's centralized people data platform and foundational datasets that power all AI verification products. Requires 10+ years experience with large-scale data pipelines, PySpark, Spark, Kafka, Iceberg, and AWS services; will mentor juniors and own architecture.
Build and operate scalable data pipelines, models, and infrastructure using Airflow, Snowflake, Databricks, AWS, and Terraform. The role requires 2+ years of data engineering experience, strong Python and SQL skills, and effective collaboration with technical and business stakeholders.
Senior Data Engineer building and owning large-scale ETL pipelines at Tatari to power reporting and measurement products from high-volume TV advertising data. Requires 5+ years experience with Python, SQL, Spark, Airflow, and strong data quality practices in a cross-functional environment.
Build and maintain internal data pipelines, metrics, dashboards, Genie spaces, and AI/agentic applications that power Databricks' ISV partner organization. Heavy hands-on Databricks user who also evaluates partner AI tools (Cursor, Claude, etc.) and feeds learnings back to product and DevRel teams.
Own end-to-end multi-million-dollar code data programs for frontier AI labs. Decompose model weaknesses, design human/synthetic/hybrid data pipelines, manage expert teams, and serve as the trusted technical partner to researchers. Requires strong coding literacy, ML familiarity, and operational ownership to deliver high-quality data under tight timelines.
OpenAI is hiring across multiple disciplines to design, build, scale, and operate its global compute infrastructure powering frontier AI models like GPT-5.6. Roles span distributed systems, ML infrastructure, hardware, manufacturing, supply chain, data center development, and physical engineering systems.
Senior Streaming Software Engineer building and operating high-throughput real-time data pipelines and streaming platforms (Kafka, Flink etc.) for observability, logs, metrics and traces across Crusoe's large-scale AI GPU cloud and data centers. Requires strong distributed systems and backend engineering experience.
Own the architecture and standards for Cribl's analytics engineering platform using dbt. Build certified models, semantic layers, and governance practices to power trusted analytics, reporting, and AI decision-making across the business.
Analytics Engineer responsible for extending core data models, designing scalable pipelines focused on market data, and owning data observability for Pave's compensation intelligence products. Requires 4+ years in data/analytics engineering, proficiency in dbt and Airflow, and experience with ML workflows in a product-facing role.
Build and maintain canonical data models, metric definitions, and dbt transformations to create a trusted, reusable data layer for internal teams and payer customers in a healthcare startup. Requires expert SQL, dbt experience, metric reconciliation, and a product-oriented approach to data quality and documentation.
Staff Software Engineer owning architecture and technical direction for scalable data integrations platform handling clinical/financial data with EMRs. Requires 10+ years experience building distributed systems, high-throughput pipelines, and deep infrastructure knowledge (GCP/K8s, databases, networking).
Build and own the facilities data pipeline for AI data center telemetry, including ingestion from industrial protocols (BACnet, Modbus, OPC UA), data quality tooling, and serving clean APIs/datasets for dashboards, controls, and ML. Requires production data pipeline experience with on-call ownership, industrial protocol integration, and full-stack debugging.