Senior/Staff Software Engineer, Data Platform
Build and own scalable streaming, storage, and distributed data-platform systems powering a unified telematics API. The role requires strong Java or JVM-language experience, production platform ownership, and expertise in distributed systems and large-scale batch or streaming data.
About the job
Responsibilities
- Design and build streaming and batch pipelines for replication, analytics, and enrichment.
- Build storage systems that scale to petabytes of time-series data while remaining fast under load.
- Improve platform primitives including data quality, lineage, stream and batch processing, and high-availability infrastructure.
- Develop AI-powered tooling that helps the platform improve itself.
- Contribute to open-source technologies when appropriate.
- Make architectural decisions that keep the platform ahead of its scale.
- Work primarily in Java and Python.
- Partner with engineering teams and customers to build reusable abstractions and components.
Requirements
- Experience building and owning production platforms or distributed systems at scale.
- Strong software engineering fundamentals and a computer science foundation.
- Hands-on experience with big-data frameworks and streaming or batch systems operating at TB-to-PB scale.
- Deep understanding of distributed systems and failure-mode design.
- Experience building durable platform systems used by other teams.
- Strong judgment in ambiguous environments and the ability to balance thorough design with rapid experimentation.
- Strong Java or another JVM language such as Scala or Kotlin, or the ability to learn it quickly.
Nice-to-haves
- Experience with data streaming and processing, including Kafka, Flink, Spark, warehouses, or lakehouses.
- Open-source contributions, particularly in data processing or storage.
- Lakehouse storage technologies such as Iceberg, Delta, Paimon, or Doris.
- Orchestration and workflow engines such as Temporal or Step Functions.
- Time-series and spatial or spatio-temporal data experience.
Technology stack
- Languages: Java, Python, TypeScript, Node.js
- Framework: Spring Boot
- Storage: AWS S3, PostgreSQL, DynamoDB, Apache Doris, Apache Iceberg, Redis
- Streaming: AWS Kinesis, Apache Kafka, Apache Flink
- Orchestration: Temporal, AWS Step Functions
- ETL: AWS Glue, Apache Spark
- AI: LangGraph, Deep Agents, Amazon Bedrock
- Infrastructure as code: Pulumi
Compensation and benefits
- Senior base salary: $200,000–$255,000 CAD, plus equity.
- Staff base salary: $235,000–$295,000 CAD, plus equity.
- Health and dental benefits with a flexible healthcare spending account.
- Personal spending account for professional development, fitness, and wellness.
- Four weeks of paid time off plus statutory holidays.
- MacBook and computer equipment.
- Downtown Toronto office and in-person work culture.
Skills
Java, Python, Scala, Kotlin, Apache Kafka, Apache Flink, Spark, Aws S3, Postgres, DynamoDB, Apache Iceberg, Redis, Temporal, Aws Kinesis, Pulumi
Similar jobs
Data Engineering jobsStaff engineer owning the analytical data layer, schema, and tiered analytics architecture. The role combines hands-on backend development with database performance optimization, observability, ingestion coordination, and measured architectural decision-making.
Leads the re-platforming of Vanta’s compliance data layer from MongoDB to schema-aware PostgreSQL across high-throughput Kafka and S3 pipelines. The role requires staff-level distributed systems expertise, migration leadership, and strong experience with relational and document data modeling.
Build and operate foundational streaming, messaging, and data pipeline infrastructure for highly scalable identity and analytics systems. The role requires 3+ years of software development experience and strengths in distributed systems, event streaming, and platform reliability.
Leads the design and improvement of large-scale data platform systems and workflows, collaborating across engineering, data science, and business teams. Requires 10+ years of relevant experience, advanced SQL and Python, cloud data tooling, and technical leadership.
Leads the migration and evolution of Vanta’s multi-tenant data platform, designing highly reliable ingestion, storage, and query systems at terabyte scale. The role requires staff-level distributed-systems expertise, Kafka and database fluency, and the ability to drive architecture across teams.