Senior Data Engineer
Build and operate scalable lakehouse infrastructure, streaming and CDC pipelines, query systems, and self-serve BI capabilities. Requires 5+ years of data engineering experience, strong Kubernetes and infrastructure-as-code expertise, and hands-on experience with distributed data platforms.
About the job
Responsibilities
- Design, build, and evolve core data platform infrastructure, including distributed query engines, orchestration, warehousing, and cataloging.
- Own lakehouse infrastructure as code, managing deployments through Terraform and Ansible on Kubernetes.
- Build and maintain low-latency streaming and CDC ingestion pipelines, plus batch ingestion paths landing in Apache Iceberg.
- Develop and scale the BI landscape to provide performant, self-serve access to lakehouse data.
- Enforce platform reliability practices, including monitoring, alerting, on-call rotations, incident response, maintenance windows, runbooks, and SLAs.
- Partner with DevOps, Analytics Engineering, and other stakeholders to close infrastructure gaps and support new data requirements.
Requirements
- 5+ years of data engineering experience, including 2+ years building and operating scalable, low-latency data platforms handling more than 100 million events per day.
- Hands-on experience running data infrastructure on Kubernetes with cloud-native tooling such as Docker and Helm.
- Production experience with infrastructure as code using Terraform, Ansible, and ArgoCD or equivalent tools.
- Deep knowledge of distributed systems, including storage, transactions, and query processing, with experience operating open-source query engines such as Trino or Presto.
- Experience with object storage and open table formats, specifically Apache Iceberg.
- Experience with streaming and CDC systems including Kafka, Redpanda, and Debezium.
- Experience with orchestration frameworks and ELT tools, including Airflow and Airbyte.
- Strong Python and SQL skills for building pipelines and platform tooling.
- Experience with Google Cloud Platform and data services such as GCS, Cloud Build, Cloud SQL, and Dataproc, or equivalent cloud services.
- Ability to work effectively in a fast-paced startup environment and adapt infrastructure to changing needs.
Nice-to-Haves
- Experience with semantic or metrics layers such as Cube, dbt, or Looker.
- Familiarity with transformation frameworks, including dbt.
- Familiarity with reverse ETL tooling such as Hightouch.
- Familiarity with data catalog and lineage tooling such as OpenMetadata or DataHub.
- Experience with data access control and governance frameworks such as Apache Ranger.
Compensation and Benefits
- Competitive salary and stock options.
- Health benefits.
- One-time USD $500 home-office setup allowance for new hires.
- USD $150 monthly stipend via a Brex Card.
Skills
Kubernetes, Docker, Helm, Terraform, Ansible, Argo CD, Trino, Apache Iceberg, Kafka, Debezium, Airflow, Python, SQL, GCP
Similar jobs
Data Engineering jobsLeads database architecture, performance, reliability, and developer-tooling initiatives for high-volume trading applications. Requires 8+ years of software engineering experience, expert MySQL skills, backend development expertise, and strong knowledge of distributed systems and database operations.
Senior Data Engineer responsible for building and operating reliable clinical and claims data pipelines, CDC systems, quality controls, and de-identified exports. The role requires 5+ years of production pipeline experience plus strong SQL, Python, Spark, and data-governance skills.
Own the company’s metric governance program by defining canonical metrics, enforcing them in semantic and catalog systems, improving data quality, and validating AI-agent outputs. Requires 5+ years in analytics or analytics engineering, strong SQL, production semantic-layer ownership, and experience with AI evaluation and data governance.
Construye y lidera la arquitectura, los pipelines y la plataforma de datos para habilitar analítica de producto, reportes financieros y experiencias self-serve. Requiere más de 7 años de experiencia, dominio de SQL, Snowflake, dbt, Python y AWS, además de experiencia con orquestación y modelado de datos.
Senior Data Engineer responsible for architecting and operating scalable data pipelines, warehouses, and analytics infrastructure. The role requires 7+ years of data or analytics engineering experience, strong SQL and modeling expertise, and proficiency with cloud, orchestration, and BI technologies.