Skip to content
ParafinParafin

Senior Software Engineer, Data Platform

Build and maintain scalable data pipelines and lakehouse infrastructure using PySpark, Databricks, and Airflow on AWS. Partner with Data Science and Engineering teams to enhance data quality, observability, and ML platform support. Requires 4+ years experience with Python, SQL, and cloud data stacks.

About the job

What You’ll Do

  • Design and build robust, highly scalable data pipelines and lakehouse infrastructure with PySpark, Databricks, and Airflow on AWS.
  • Improve the data platform development experience for Engineering, Data Science, and Product by creating intuitive abstractions, self‑service tooling, and clear documentation.
  • Own and maintain core data pipelines and models that power internal dashboards, ML models, and customer-facing products.
  • Own the Data & ML platform infrastructure using Terraform, including end‑to‑end administration of Databricks workspaces: manage user access, monitor performance, optimize configurations (e.g., clusters, lakehouse settings), and ensure high availability of data pipelines.
  • Lead projects to improve data quality, testing, observability, and cost efficiency across existing pipelines and backend systems (e.g., migrating Databricks SQL pipelines to dbt, scaling data ingestion, improving data-lineage tracking, and enhancing monitoring).
  • Act as the primary engineering partner for the Data Science team—embedded closely to gather requirements, design scalable solutions, and provide end-to-end support on all engineering aspects of their work.
  • Work closely with backend engineers and data scientists to design performant data models and support new product development initiatives.
  • Share best practices and mentor other engineers working on data-centric systems.

What We’re Looking For

  • 4+ years of experience in software engineering with a strong background in data infrastructure, pipelines, and distributed systems.
  • Advanced proficiency in Python and SQL.
  • Hands-on Spark development experience.
  • Expertise with modern cloud data stacks—AWS (S3, RDS), Databricks, and Airflow—and lakehouse architectures.
  • Hands‑on experience with foundational data‑infrastructure technologies such as Hadoop, Hive, Kafka (or similar streaming platforms), Delta Lake/Iceberg, and distributed query engines like Trino/Presto.
  • Familiarity with ingestion frameworks, developer‑experience tooling, and best practices for data versioning, lineage, partitioning, and clustering.
  • Strong problem-solving skills and a proactive attitude toward ownership and platform health.
  • Excellent communication and collaboration skills, especially in cross-functional settings.

Bonus Points

  • Experience with AWS infrastructure using Terraform.
  • Familiarity with observability tools (e.g., Datadog) and cost tracking in cloud environments.
  • Experience with financial systems or building platforms in a fintech setting.
  • Prior work on ML infrastructure: Feature stores (e.g., Tecton), ML model lifecycle (training, deployment, monitoring, retraining), real-time inference.
  • Contributions to internal tooling or open-source projects in the data ecosystem.

What We Offer

Salary Range: $230k-$265k

  • Equity grant
  • Medical, dental & vision insurance
  • Work from home flexibility
  • Unlimited PTO
  • Commuter benefits
  • Free lunches
  • Paid parental leave
  • 401(k)
  • Employee assistance program

Skills

Python, SQL, Pyspark, Databricks, Airflow, AWS, Terraform, Spark, Kafka, Delta Lake, dbt, Trino, Hadoop, Hive, Datadog

Zoox

Zoox

Foster City, CA

Lead Data Platform Engineer - Enterprise, Data & AI
$230k+/yrHybrid10+ YOEData Engineering

Leads the architecture, scaling, security, governance, and cost optimization of enterprise and AI data platforms. Requires 10+ years of data or software engineering experience, with expertise in production data foundations, CI/CD, governance, security, and performance optimization.

Garner Health

Garner Health

New York, NY

Senior Data Engineer
$220k+/yrHybrid5+ YOEData Engineering

Build and scale data pipelines, reusable datasets, and validation frameworks supporting business intelligence, marketing, and data science. The role requires strong Python and SQL skills, modern data-stack experience, and at least four years of software or data engineering experience.

Prompt Health

Prompt Health

United States

Senior Database Reliability Engineer
$220k+/yrRemote6+ YOEData Engineering

Own the reliability, performance, observability, scalability, and cost efficiency of large Aurora MySQL production environments supporting healthcare applications. The role requires 6+ years of database engineering experience, deep MySQL and AWS expertise, and strong skills in automation, incident response, and query optimization.

OnePay

OnePay

United States

Analytics Engineering Manager
$220k+/yrRemote7+ YOEData Engineering

Leads an analytics engineering team that transforms raw data into reliable, actionable insights for product, marketing, and operations. The role requires 7+ years in data or analytics engineering, management experience, and advanced SQL, Databricks, and dbt expertise.

Decagon

Decagon

San Francisco, CA
Senior Software Engineer, Data Infrastructure
$200k+/yrOn-site5+ YOEData Engineering

Senior Data Infrastructure Engineer responsible for building and operating reliable, low-latency streaming and batch data systems that support AI products. Requires 5+ years of production data infrastructure experience and expertise with technologies such as Kafka, Flink, ClickHouse, and Terraform.