Software Engineer III
Build and operate large-scale backend data pipelines and ETL/ELT workflows for acquiring, validating, and storing high-volume business data using Java, Spark, Airflow, Beam, Kafka and GCP services. Requires 3+ years experience in production data engineering.
About the job
What You'll Do
- Design and build backend data pipelines that ingest, validate, normalize, enrich, and store high-volume raw data from multiple sources.
- Develop ETL/ELT workflows and processing jobs using technologies such as Apache Airflow, Apache Beam, Google Dataflow, DataProc, Spark, Kafka, or Pub/Sub.
- Implement new features in Java-based services and data processing applications that support data acquisition at scale.
- Work with batch and streaming architectures for scheduled, near-real-time, and event-driven data flows.
- Improve data quality through schema validation, deduplication, enrichment, monitoring, retries, and controlled backfills.
- Contribute to observability for pipeline health, throughput, latency, cost, and error rates.
- Collaborate with product managers, data teams, and platform teams to translate business requirements into reliable technical solutions.
- Help design, plan, and execute the roadmap for next-generation data acquisition technologies.
Must-Have Qualifications
Data Engineering and Pipelines
- 3+ years of professional software engineering experience.
- Solid experience building and operating production data pipelines, ETL/ELT workflows, or data processing systems.
- Strong proficiency with Java and object-oriented programming.
- Hands-on experience with data processing and orchestration technologies such as Apache Beam, Apache Airflow, Spark, Google Dataflow, or DataProc.
- Experience with streaming technologies such as Kafka, Google Pub/Sub, or similar systems.
- Understanding of batch processing, streaming processing, data modeling, schema design, and data quality practices.
Backend, Cloud, and Operations
- Experience building backend services, APIs, or distributed systems that run in production.
- Experience with at least one cloud provider, preferably GCP.
- Familiarity with cloud data and compute services such as BigQuery, GCS, GKE, Dataflow, DataProc, and Pub/Sub.
- Practical knowledge of SQL, large-scale storage/query systems, logging, monitoring, and alerting.
- Ability to troubleshoot pipeline failures, data issues, performance bottlenecks, and operational incidents.
General
- Bachelor's degree in Computer Science, Software Engineering, or a related field.
- Strong problem-solving skills and attention to detail.
- Pragmatic engineering judgment with the ability to balance quality, speed, and business impact.
Nice to Have
- Experience with Kubernetes, especially GKE or EKS, for distributed workloads.
- Experience with Snowflake, BigQuery, Starburst/Trino, or similar data warehouses and query engines.
- Experience with Terraform or other infrastructure-as-code tools.
- Exposure to data integration patterns involving CRM systems, email/calendar APIs, third-party feeds, or large external datasets.
- Experience in a B2B data company, data marketplace, or data-as-a-product environment.
Skills
Java, Apache Airflow, Apache Beam, Google Dataflow, Dataproc, Spark, Kafka, Pub/Sub, GCP, BigQuery, Gcs, GKE, SQL, Kubernetes, Terraform
Similar jobs
Data Engineering jobsBuild and scale AWS-based data infrastructure, pipelines, knowledge graphs, and APIs for a CTV performance advertising platform. The role requires production data engineering experience with Spark, Scala, AWS, SQL, and large-scale services, plus a bachelor's degree.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.
Own the data platform infrastructure supporting Mercor’s data-driven teams, including Snowflake governance, access controls, ingestion, orchestration, dashboard governance, and cost management. The role requires strong Snowflake administration, SQL, Python, production pipeline ownership, and CDC/streaming experience.
Build and own Stuut’s foundational data platform, including ingestion pipelines, canonical models, semantic layers, and observability. The role requires 3+ years of production data pipeline experience with Python, SQL, cloud warehouses, and ETL/ELT tooling.
Build scalable analytics engineering infrastructure, SaaS data models, and AI-enabled workflows that support enterprise decision-making. The role requires 3–6 years of hands-on analytics or data engineering experience, strong SQL and modern data modeling expertise, and cloud data warehouse experience.