Skip to content
BaselayerBaselayer

Data Engineer

Build and operate production data pipelines and transformation layers that turn heterogeneous business, identity, and fraud data into reliable inputs for entity resolution, scoring, and customer APIs. The role requires at least one year of data engineering experience with Python, SQL, cloud platforms, and modern pipeline tooling.

About the job

Responsibilities

  • Build and maintain ETL/ELT pipelines ingesting and normalizing public records, web signals, and fraud telemetry.
  • Develop data models and transformation layers using tools such as Dataflow, Spark, and Airflow to support fraud detection, KYB, and customer-facing APIs.
  • Implement data quality checks, observability tooling, and alerting.
  • Tune pipelines and queries for performance, freshness, and cost in the cloud data warehouse.
  • Collaborate with data scientists, ML engineers, and product teams to provide well-modeled data for entity resolution and scoring.
  • Help ensure pipelines meet security and regulatory standards for sensitive data, including SOC 2, GDPR, and KYC/KYB.
  • Document systems and communicate with technical and non-technical stakeholders.

Requirements

  • 1+ years of experience in data engineering with Python, SQL, and cloud-native data platforms.
  • Experience building and maintaining production ETL/ELT pipelines.
  • Working knowledge of modern data-stack tooling such as Dataflow, Spark, or Airflow.
  • Hands-on experience with cloud data warehouses or data lakes, such as BigQuery or Snowflake.
  • Strong data-modeling fundamentals and focus on data integrity and reliability.
  • Comfort working with structured and unstructured data.

Nice-to-Haves

  • Interest in AI/ML infrastructure.
  • Experience with streaming or real-time data systems such as Kafka or Pub/Sub.
  • Exposure to KYC/KYB, fraud, risk, or underwriting data.
  • GCP experience, including BigQuery, Cloud Run, Dataflow, or Pub/Sub.
  • Experience working in an early-stage environment.

Compensation and Benefits

  • $120,000–$150,000 salary plus equity.
  • Flexible PTO.
  • 100% employer-paid health, dental, and vision premiums.
  • 401(k) with company match.
  • HSA contributions on applicable plans.
  • $250 monthly gym stipend.

Skills

Python, SQL, Dataflow, Spark, Apache Airflow, ETL, ELT, BigQuery, Snowflake, Kafka, GCP, Pub/Sub, Data Modeling, Data Quality, Cloud Run

TheGuarantors

TheGuarantors

New York, NY

Data Engineer
$115k+/yrHybrid1+ YOEData Engineering

Build and maintain reliable data pipelines, warehouses, and lightweight data applications supporting analytics, operational models, and AI/ML workflows. The role requires SQL, Python, ETL, and data warehousing knowledge, with 1+ year of relevant experience preferred.

Seeq

Seeq

United States

AI Customer Insights Engineer
$115k+/yrRemoteData Engineering

Builds customer intelligence workflows that turn product usage, engagement, and contract data into actionable insights for Customer Success. The role requires strong Python and SQL skills, a bachelor's degree, and an interest in applied AI, automation, and predictive customer health modeling.

Vercel

Vercel

San Francisco, CA
Software Engineer, Data Platform
$130k+/yrHybrid1+ YOEData Engineering

Build and operate the streaming, storage, query, and self-service infrastructure underlying the company’s data platform. The role suits an early-career engineer with 1–3 years of experience, a computer science bachelor’s degree, programming skills, and interest in distributed systems.

Astera

Astera

Washington, DC

Data and Partnerships Associate
$135k+/yrOn-site5+ YOEData Engineering

Manages the full lifecycle of scientific research data, including governance, metadata, repositories, open-science publishing, and AI/ML compatibility. The role also builds partnerships with federal research organizations and requires a bachelor’s degree plus 5–7+ years of relevant experience.

Garner Health

Garner Health

New York, NY

Data Engineer III
$166k+/yrHybrid2+ YOEData Engineering

Build and optimize scalable data pipelines, reusable datasets, and federated data quality systems for healthcare analytics. The role requires at least 2 years of data or software engineering experience and strong Python, SQL, AWS, orchestration, database, and warehouse expertise.