Staff Data Engineer, tvScientific
Lead design and implementation of scalable identity resolution and data governance platforms. Build pipelines for identity data management, ensure privacy compliance, and partner with teams to deliver reliable data services. Requires 5+ years Spark/Scala experience.
About the job
What you'll do:
Identity Services:
- Design and maintain a scalable identity resolution platform
- Build pipelines and services to ingest, normalize, link, and version identity data across multiple sources
- Ensure deterministic and probabilistic matching logic that is transparent, auditable, and measurable
- Partner with product and analytics teams to expose identity data through reliable, well-documented APIs and datasets
- Build and operate batch and streaming pipelines using modern data stack tools
- Create clear documentation, standards, and runbooks for identity and governance systems
Data Governance & Trust:
- Own data governance foundations including data lineage, quality checks, schema enforcement, and access controls
- Implement privacy-by-design principles (PII handling, consent enforcement, retention policies)
- Collaborate with legal, privacy, and security teams to operationalize regulatory requirements (e.g., GDPR, CCPA)
- Establish monitoring and alerting for data quality, freshness, and integrity
What we're looking for:
- Data engineering experience with proven track record building data infrastructure using Spark with Scala
- Proven experience building data infrastructure using Spark with Scala for at least 5 years
- Experience in delivering significant technical initiatives and building reliable, large scale services
- Experience in delivering APIs backed by relationship-heavy datasets
- Experience implementing data governance practices, including data quality, metadata management, and access controls
- Strong understanding of privacy-by-design principles and handling of sensitive or regulated data
- Familiarity with data lakes, cloud warehouses, and storage formats
- Strong proficiency in AWS services
- Successful design and implementation of scalable and efficient data infrastructure
- High attention to detail in implementation of automated data quality checks
- Effective collaboration with cross-functional teams
- Excellent written and verbal communication skills
- Bachelor's degree in Computer Science or a related field
Skills
Spark, Scala, AWS, Identity Resolution, Data Pipelines, Data Governance, Data Lakes, Streaming Pipelines, APIs, Data Quality
Similar jobs
Data Engineering jobsLeads the architecture and hands-on development of a knowledge-graph-centered data platform for AI and autonomy workflows. The role requires deep distributed data systems expertise, strong Go or Python engineering skills, and experience with storage, APIs, infrastructure, and production reliability.
Build and maintain data pipelines, analytics models, dashboards, and external data products while partnering with engineering, product, implementation teams, and customers. The role requires 5+ years of analytics or data engineering experience, strong dbt and SQL expertise, and customer-facing collaboration skills.
Architects scalable data systems and platforms using distributed technologies like Spark, Kafka, and AWS. Mentors engineers and drives innovation on large-scale data projects, requiring 8+ years experience and expertise in data infrastructure.
Leads architecture and technical governance for an enterprise-scale data platform, driving distributed systems, data modeling, cloud infrastructure, and operational best practices. Requires 7+ years of engineering experience and a bachelor’s degree.
Leads the design, operation, and technical direction of Pinterest’s data workflow and context control planes, driving reliability, scalability, AI-native capabilities, and open-source contributions. Requires 10+ years of distributed-systems experience, infrastructure expertise, and proficiency in Python or Java.