Skip to content
ZoomInfoZoomInfo

Senior Software Engineer - Crawler

Builds and operates large-scale web crawling, extraction, and data engineering infrastructure processing billions of pages. The role requires 5+ years of software engineering experience, strong distributed-systems fundamentals, and proficiency with Java or Python, cloud platforms, Kubernetes, and ETL technologies.

About the job

Responsibilities

  • Design and implement components of scalable, fault-tolerant web crawling and extraction pipelines.
  • Write clean, production-grade code in Java and Python.
  • Build and operate ETL/ELT pipelines for large-scale data extraction and transformation.
  • Work with cloud infrastructure on GCP and AWS, primarily on GKE.
  • Improve observability, reliability, and operational excellence across contributed systems.
  • Partner with product and data science teams to deliver impactful solutions.
  • Contribute to code reviews, documentation, and knowledge sharing.
  • Stay current with web technologies, anti-crawling mechanisms, and AI-powered extraction approaches.

Requirements

  • 5+ years of professional software engineering experience building production systems.
  • Strong computer science fundamentals, including algorithms, data structures, concurrency, and distributed systems.
  • Proficiency in Java and/or Python.
  • Experience owning features end-to-end, from design through deployment and operation.
  • Ability to make sound component-level architectural decisions.
  • Hands-on experience with cloud data warehouses such as BigQuery or Snowflake.
  • Experience designing and operating large-scale ETL/ELT pipelines.
  • Experience with orchestration tools such as Apache Airflow.
  • Experience with streaming or event-driven systems such as Apache Kafka.
  • Production experience on GCP or AWS; multi-cloud exposure is a plus.
  • Hands-on experience with Kubernetes, including GKE or EKS, for distributed workloads.
  • Familiarity with infrastructure-as-code tools such as Terraform.
  • Strong communication, pragmatic problem-solving, and ability to operate in ambiguity.

Nice-to-Haves

  • Web crawling at scale using Scrapy or similar frameworks.
  • Proxy infrastructure, rotation strategies, or anti-bot evasion techniques.
  • Structured and unstructured web data extraction from diverse site architectures.
  • SERP extraction.
  • AI/LLM-based extraction approaches applied to HTML at scale.
  • Experience in a B2B data company or data-as-a-product environment.

Compensation and Benefits

  • US base salary: $140,000–$220,000 USD.
  • Additional compensation such as bonus, commission, equity, and other benefits may apply.
  • Comprehensive benefits and holistic mind, body, and lifestyle programs.

Skills

Java, Python, GCP, AWS, Kubernetes, GKE, BigQuery, Snowflake, Apache Airflow, Apache Kafka, Terraform, Scrapy, ETL, Distributed Systems, Web Crawling

Mark43

Mark43

Boston, MA

Database Admin Engineer
$140k+/yrOn-site7+ YOEData Engineering

Own and scale Mark43’s production database infrastructure, focusing on MySQL administration, monitoring, performance, reliability, upgrades, and recovery. The role requires at least seven years of production database experience, cloud infrastructure expertise, and familiarity with Terraform.

ZoomInfo

ZoomInfo

Waltham, MA

Senior Software Engineer
$140k+/yrHybrid5+ YOEData Engineering

Build and operate large-scale data acquisition pipelines, distributed processing systems, and production Java services across batch and streaming workloads. The role requires 5+ years of backend or data engineering experience, strong distributed-systems expertise, and proficiency with cloud data technologies.

Square

Square

California

People Intelligence Architect
$143k+/yrOn-site8+ YOEData Engineering

Senior individual contributor who architects and builds end-to-end people data systems, predictive models, and AI-agent workflows. Requires 8+ years of experience across data engineering and data science, with expertise in Python, SQL, machine learning, sensitive HR data, and agentic AI.

NinjaTrader

NinjaTrader

Chicago, IL

Marketing Analytics Engineer
$135k+/yrHybrid6+ YOEData Engineering

Senior individual contributor responsible for architecting shared dbt models, marketing attribution, data quality, and AI-driven analytics workflows. Requires 6+ years in analytics engineering or data, deep dbt and SQL expertise, modern data-stack experience, and strong marketing measurement knowledge.

Wrapbook

Wrapbook

United States
Senior Analytics Engineer II
CA$148k+/yrRemote5+ YOEData Engineering

Build and own Wrapbook’s analytics layer, including production pipelines, governed data models, canonical datasets, self-serve analytics, and monitoring. The role requires strong SQL and Python, modern warehouse experience, and 4+ years in data or analytics engineering.