Staff Data Engineer
Founding Data Engineer to architect Payabli's data platform from scratch: design lakehouse/warehouse, build pipelines, model financial data, and establish governance for a regulated fintech environment.
About the job
What You'll Do
- Architect the platform. Set our warehouse/lakehouse direction and stand up the data lake and layered architecture that turns our raw system of record into trustworthy, queryable, intelligence-ready data.
- Build the pipelines. Design and run batch and streaming pipelines that move data reliably out of our production systems - CDC, ELT, and real-time where it matters.
- Model the data. Define the canonical datasets and models the whole company depends on, getting the grain, semantics, and contracts right.
- Own reliability and accuracy. This is financial data, so correctness is non-negotiable. You'll own data quality, observability, integrity checks, and the testing and monitoring that let us trust it.
- Build for a regulated environment. Design in role-based access, masking, lineage, and auditability from day one, and keep sensitive financial data out of places it doesn't belong.
- Enable AI/ML and analytics. Build the feature pipelines and trustworthy data foundation our intelligence work relies on, moving us from systems of record toward systems of intelligence and action.
- Set the standard. Establish the practices, tooling, and CI/CD for data that the future team inherits.
What We're Looking For
- 8+ years building production data systems, with a track record of owning architecture and seeing big decisions through to production.
- Expert SQL and strong Python.
- Deep experience in at least one modern lakehouse/warehouse ecosystem (e.g., Snowflake with dbt and Fivetran, or Databricks with Spark, Delta Lake, and Unity Catalog).
- Strong data modeling skills (dimensional, normalized, or Data Vault).
- Experience with pipeline orchestration (Airflow, Dagster, Prefect, or equivalent) and large-scale processing (such as Spark).
- Production experience on a major cloud (AWS, GCP, or Azure), including security and cost patterns.
- Experience working with sensitive or regulated data (access controls, encryption, governance).
Nice to Haves
- Payments, fintech, or other regulated-domain experience, including familiarity with PCI DSS and tokenization/vaulting patterns.
- Streaming infrastructure (Kafka, Kinesis, Flink).
- Data governance, lineage, and observability tooling (Unity Catalog, Snowflake Horizon, Monte Carlo, Great Expectations, OpenLineage).
- Experience supporting ML/AI workloads (feature stores, training/inference pipelines, MLflow).
- Interest in growing into people leadership.
Compensation & Benefits
- Competitive salary
- Stock options with the potential to unlock more equity as we grow
- Flexible PTO and paid parental leave
- Medical, dental, & vision insurance
- 401K, HSA, pre-tax savings programs
Skills
SQL, Python, Snowflake, dbt, Fivetran, Databricks, Spark, Delta Lake, Unity Catalog, Airflow, Dagster, Prefect, AWS, GCP, Azure
Similar jobs
Data Engineering jobsLeads enterprise data engineering strategy, architecture, delivery, governance, and technical leadership across the organization. Requires extensive data engineering experience, advanced data modeling and warehouse expertise, and strong PySpark, SQL, and Python skills.
Staff Software Engineer leading development of data catalog and metadata infrastructure for discovery, governance, lineage, and quality across Airbnb’s data ecosystem. Requires 9+ years of software engineering experience focused on data infrastructure and strong programming and distributed data technology skills.
Staff Software Engineer responsible for designing and scaling Ray Data’s distributed data-processing infrastructure for large-scale AI training and inference. Requires 6+ years of production software and architectural ownership experience, plus deep distributed-systems expertise and strong Python skills.
Leads the design, operation, and technical direction of Pinterest’s data workflow and context control planes, driving reliability, scalability, AI-native capabilities, and open-source contributions. Requires 10+ years of distributed-systems experience, infrastructure expertise, and proficiency in Python or Java.
Build and operate foundational streaming, messaging, and data pipeline infrastructure for highly scalable identity and analytics systems. The role requires 3+ years of software development experience and strengths in distributed systems, event streaming, and platform reliability.