Skip to content

Principal Java Data Engineer

Designs, develops, and maintains large-scale data platforms and pipelines using Java microservices. Leads technical direction, mentors engineers, and ensures data quality, governance, and observability in cloud-native environments. Requires 10+ years experience with 4+ in data pipelines.

About the job

Principal Java Data Engineer

Role Summary

Contribute to all phases of the software development life cycle, designing, developing, and maintaining large-scale Data Platform and data pipelines based on microservices architecture. Hands-on leadership role enhancing batch and real-time data solutions, mentoring team members, and delivering business and technical objectives.

Day-to-Day Responsibilities

  • Lead and guide the design and implementation of scalable distributed systems based on Java microservices
  • Engineer and optimize data pipelines using solutions like Apache Hudi, Apache Trino, Azure ADLS
  • Collaborate cross-functionally with product, analytics, and AI teams to ensure data is a strategic asset
  • Advance ongoing modernization efforts, deepening adoption of event-driven architectures and cloud-native technologies
  • Drive adoption of best practices in data governance, observability, and performance tuning for data workloads
  • Embed data quality in processing pipelines by defining schema contracts, implementing transformation tests and data assertions, enforcing backward-compatible schema evolution, and automating checks for freshness, completeness, and accuracy across batch and streaming paths before production deployment
  • Establish robust observability for data pipelines by implementing metrics, logging, and distributed tracing for streaming jobs, defining SLAs and SLOs for latency and throughput, and integrating alerting and dashboards to enable proactive monitoring and rapid incident response
  • Foster a culture of quality through peer reviews, providing constructive feedback and seeking input on your own work

Qualifications

  • Principal Software Data Engineer with at least 10 years of professional experience in software or data engineering, including a minimum of 4 years focused on data pipelines (batch and streaming)
  • Proven experience driving technical direction and mentoring engineers while delivering complex, high-scale solutions as a hands-on contributor
  • Strong understanding of event-driven architectures and distributed systems, with hands-on experience implementing resilient, low-latency pipelines
  • Practical experience with cloud platforms (AWS, Azure, or GCP) and containerized deployments for data workloads
  • Fluency in data quality practices and CI/CD integration, including schema management, automated testing, and validation frameworks (e.g., dbt, Great Expectations)
  • Operational excellence in observability, with experience implementing metrics, logging, tracing, and alerting for data pipelines using modern tools
  • Solid foundation in data governance and performance optimization, ensuring reliability and scalability across batch and streaming environments
  • Proven experience with Lakehouse architectures and related technologies, including Apache Hudi, Azure ADLS Gen2, HDFS, and other big data technologies (Trino, Databricks, Spark)
  • Strong collaboration and communication skills, with the ability to influence stakeholders and evangelize modern data practices within your team and organization

Skills

Java, Apache Hudi, Apache Trino, Azure Adls, Spark, Databricks, dbt, Great Expectations, Kubernetes, Event-Driven Architectures

ZoomInfo

ZoomInfo

United States

Senior Principal Software Engineer
$164k+/yrRemote10+ YOEData Engineering

Architects and builds ZoomInfo’s distributed data-platform infrastructure, including federated GraphQL access, real-time pipelines, indexing, and observability. The role requires 10+ years of software engineering experience, cloud-native expertise, and strong distributed-systems design skills.

Pinterest

Pinterest

United States

Principal Engineer, Big Data Platform
$243k+/yrRemote7+ YOEData Engineering

Leads the strategy, architecture, and scaling of Pinterest’s big data and AI infrastructure across petabyte-scale workloads. Requires principal-level technical leadership, extensive Kubernetes or big data platform experience, and proficiency in modern data and cloud technologies.

Orum

Orum

Austin, TX

Staff / Principal Data Engineer
No salary listedHybrid8+ YOEData Engineering

Own the architecture, reliability, and evolution of a modern data platform spanning data engineering and analytics/BI. The role requires 8+ years of experience plus deep expertise in SQL, dbt, ClickHouse, BigQuery, Looker, streaming systems, and distributed data platforms.

Octus

Octus

United States

Principal Data Engineer
No salary listedRemote8+ YOEData Engineering

Leads the strategy, architecture, and hands-on development of scalable AWS data ingestion and transformation platforms. Requires expert Python and SQL skills, Terraform and cloud-native pipeline experience, and 8+ years in data engineering or backend development, with technical leadership responsibilities.

IntusCare

IntusCare

United States

Staff Software Engineer, Data
$180k+/yrRemote7+ YOEData Engineering

Staff-level engineer responsible for the technical direction, reliability, and evolution of a cloud ELT platform supporting healthcare data products. The role requires 7+ years of software or data engineering experience, deep SQL/Python and modern data-platform expertise, and strong architectural and mentoring leadership.