Build and operate large-scale backend data pipelines and ETL/ELT workflows for acquiring, validating, and storing high-volume business data using Java, Spark, Airflow, Beam, Kafka and GCP services. Requires 3+ years experience in production data engineering.
112k – 176k/yr
Hybrid3+ YOEData Engineering
About the role
What You'll Do
Design and build backend data pipelines that ingest, validate, normalize, enrich, and store high-volume raw data from multiple sources.
Develop ETL/ELT workflows and processing jobs using technologies such as Apache Airflow, Apache Beam, Google Dataflow, DataProc, Spark, Kafka, or Pub/Sub.
Implement new features in Java-based services and data processing applications that support data acquisition at scale.
Work with batch and streaming architectures for scheduled, near-real-time, and event-driven data flows.
Improve data quality through schema validation, deduplication, enrichment, monitoring, retries, and controlled backfills.
Contribute to observability for pipeline health, throughput, latency, cost, and error rates.
Collaborate with product managers, data teams, and platform teams to translate business requirements into reliable technical solutions.
Help design, plan, and execute the roadmap for next-generation data acquisition technologies.
Must-Have Qualifications
Data Engineering and Pipelines
3+ years of professional software engineering experience.
Solid experience building and operating production data pipelines, ETL/ELT workflows, or data processing systems.
Strong proficiency with Java and object-oriented programming.
Hands-on experience with data processing and orchestration technologies such as Apache Beam, Apache Airflow, Spark, Google Dataflow, or DataProc.
Experience with streaming technologies such as Kafka, Google Pub/Sub, or similar systems.
Understanding of batch processing, streaming processing, data modeling, schema design, and data quality practices.
Backend, Cloud, and Operations
Experience building backend services, APIs, or distributed systems that run in production.
Experience with at least one cloud provider, preferably GCP.
Familiarity with cloud data and compute services such as BigQuery, GCS, GKE, Dataflow, DataProc, and Pub/Sub.
Practical knowledge of SQL, large-scale storage/query systems, logging, monitoring, and alerting.
Ability to troubleshoot pipeline failures, data issues, performance bottlenecks, and operational incidents.
General
Bachelor's degree in Computer Science, Software Engineering, or a related field.
Strong problem-solving skills and attention to detail.
Pragmatic engineering judgment with the ability to balance quality, speed, and business impact.
Nice to Have
Experience with Kubernetes, especially GKE or EKS, for distributed workloads.
Experience with Snowflake, BigQuery, Starburst/Trino, or similar data warehouses and query engines.
Experience with Terraform or other infrastructure-as-code tools.
Exposure to data integration patterns involving CRM systems, email/calendar APIs, third-party feeds, or large external datasets.
Experience in a B2B data company, data marketplace, or data-as-a-product environment.
Builds and optimizes big data pipelines processing billions of signals for technographic data services. Requires 5+ years experience with Spark, Hadoop, Java/Scala, and distributed systems; collaborates with data scientists on ML integration.
112k – 176k/yr
Hybrid5+ YOEData Engineering
Founding Analytics Engineer
Joyful HealthNew York, NY
Founding Analytics Engineer builds analytics infrastructure from scratch, designs data models and metrics frameworks for healthcare financial data, and creates self-service dashboards. Requires 5+ years experience with SQL, dbt, Python, and cloud data warehouses; NYC-based hybrid role.
120k – 235k/yr
Hybrid5+ YOEData Engineering
Software Engineer, Reconciliation & Reporting
SquareCalifornia
Build and maintain data pipelines for reconciling card-network settlements against internal transaction data to power accurate financial, tax, and regulatory reporting at scale for a global payments platform. Requires 3+ years experience with batch/event-driven pipelines, orchestration tools, and strong SQL/programming skills in a high-stakes financial environment.
121k – 213k/yr
On-site3+ YOEData Engineering
Data Engineer
BrexSan Francisco, CA
Build and maintain scalable data models, pipelines, and Core Data tables to transform raw data into actionable insights. Collaborate with data scientists and business teams using SQL, DBT, Snowflake, Airflow, and Python.
121k – 151k/yr
Hybrid3+ YOEData Engineering
Partner Engineer: Partner Intelligence, AI & Apps
DatabricksSan Francisco, CA
Build and maintain internal data pipelines, metrics, dashboards, Genie spaces, and AI/agentic applications that power Databricks' ISV partner organization. Heavy hands-on Databricks user who also evaluates partner AI tools (Cursor, Claude, etc.) and feeds learnings back to product and DevRel teams.