Infrastructure Engineer
Build and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.
About the job
Responsibilities
- Design, build, and operate scalable data infrastructure for high-volume ingestion, transformation, storage, and access.
- Architect and implement a secure, reliable, and governed partner data-sharing platform.
- Build foundational components of an identity graph connecting customer, account, transaction, partner, and behavioral data.
- Improve data quality, lineage, observability, reliability, governance, and access controls across pipelines and datasets.
- Partner with Data, Risk, Product, Engineering, Security, and Compliance teams to define platform capabilities.
- Build reusable abstractions, tooling, frameworks, APIs, and documentation for safe data creation, discovery, consumption, and sharing.
- Support batch and real-time data use cases while balancing latency, correctness, cost, and operational complexity.
- Address entity resolution, data freshness, schema evolution, privacy boundaries, and partner-specific data contracts.
- Improve systems through automation, monitoring, alerting, incident response, and capacity planning.
- Promote reproducible pipelines, clear ownership, robust testing, and operational excellence.
Requirements
- 5+ years of hands-on engineering experience building data platforms, distributed systems, infrastructure, or backend systems at scale.
- Experience designing and operating production data pipelines with modern processing frameworks and orchestration tools.
- Deep understanding of data modeling, data quality, schema management, lineage, observability, and governance.
- Experience with cloud data infrastructure, including object storage, warehouses, lakehouse architectures, streaming systems, and compute platforms.
- Ability to design systems meeting complex data access, privacy, security, and compliance requirements.
- Strong backend engineering fundamentals and proficiency in one or more of Python, Java, Scala, Go, or Kotlin.
- Experience building platforms or services for multiple internal customers.
- Strong systems thinking and ability to evaluate latency, consistency, reliability, scalability, cost, and developer experience trade-offs.
- Excellent analytical, problem-solving, communication, ownership, and collaboration skills.
- Bachelor’s degree in Computer Science, Computer Engineering, Information Systems, or equivalent practical experience.
Nice to Have
- Partner-facing or externally shared data platform experience, including access control, auditing, and contract management.
- Experience with identity resolution, entity matching, graph-based data systems, customer 360 platforms, or knowledge graphs.
- Experience in regulated environments such as fintech, banking, lending, payments, SOC, or PCI.
- Experience with Kafka, Flink, Spark Streaming, or similar real-time technologies.
- Experience with Snowflake, Databricks, dbt, Airflow, Dagster, Iceberg, Delta Lake, or similar modern data-stack tools.
- Familiarity with privacy-preserving data sharing, clean rooms, data contracts, consent management, or fine-grained authorization.
- Experience with observability, automated validation, backfills, incident response, and service-level objectives.
Compensation and Benefits
- $170,000–$230,000 annual compensation, plus equity.
- Flexible paid time off.
- Fully covered healthcare, including dependent coverage.
- Access to One Medical and an FSA.
- 20 weeks of paid parental leave for primary caregivers and 8 weeks for all new parents.
- Choice of configured work computer.
Skills
Python, Java, Scala, Go, Kotlin, Apache Kafka, Apache Flink, Spark, Snowflake, Databricks, dbt, Apache Airflow, Dagster, Apache Iceberg, Delta Lake
Similar jobs
Data Engineering jobsOwn the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build and operate reliable, production-grade data pipelines, warehouse infrastructure, and trusted datasets supporting company-wide analytics and AI initiatives. The role requires 3+ years of production data engineering experience, strong SQL and Python skills, and experience with Snowflake, dbt, cloud infrastructure, and orchestration.
Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.