# Staff Software Reliability Engineer - Data Platform

**Company:** [Okta](https://hotfix.jobs/companies/okta)
**Location:** Toronto, Canada
**Role:** DevOps / SRE
**Salary:** CA$160k – CA$220k/yr
**Experience:** 7+ years
**Skills:** Java, Kinesis, Kafka, Flink, Spark, Apache Beam, Snowflake, Databricks, Hadoop, Kubernetes, Mesos, AWS, Distributed Systems, Multithreading, Incident Management
**Posted:** 2026-07-31

> The Staff Software Reliability Engineer will design, build, optimize, and operate scalable streaming and distributed data-platform infrastructure supporting analytics and machine learning. The role requires 5+ years of industry experience, strong software engineering skills, and expertise with distributed data technologies and reliability practices.

## Job Description

## Responsibilities
- Design, implement, and own data-intensive, high-performance, scalable platform components.
- Collaborate with engineering teams, architects, and cross-functional partners on project development, design, and implementation.
- Conduct and participate in design reviews, code reviews, analysis, and performance tuning.
- Coach and mentor engineers.
- Debug production issues across services and multiple levels of the stack.
- Participate in the on-call rotation and incident management.

## Requirements
- 5+ years of industry experience.
- 2+ years of experience with an object-oriented language, preferably Java.
- Hands-on experience with cloud-based distributed computing technologies, including:
  - Messaging systems such as Kinesis and Kafka.
  - Data processing systems such as Flink, Spark, and Beam.
  - Storage and compute systems such as Snowflake, Databricks, and Hadoop.
  - Coordinators and schedulers such as those in Kubernetes, Hadoop, and Mesos.
- Experience developing and tuning highly scalable distributed systems.
- Excellent grasp of software engineering principles.
- Solid understanding of multithreading, garbage collection, and memory management.
- Experience with reliability engineering, particularly data quality, data observability, and incident management.

## Nice to Have
- Experience maintaining security, encryption, identity management, or authentication infrastructure.
- Experience using major public cloud providers to build mission-critical, high-volume services.
- Experience developing data integration applications for large-scale, petabyte-scale environments across batch and online systems.
- Contributions to distributed systems or experience using high-volume or critical systems such as Kafka or Hadoop.
- Experience developing Kubernetes-based services on AWS.

## Compensation and Benefits
- Annual base salary for candidates located in Canada: **$160,000–$220,000 CAD**.
- Equity, where applicable, bonus, and benefits including health, dental, and vision insurance, RRSP matching, healthcare spending, telemedicine, and paid leave including PTO and parental leave.

## Similar jobs

- [Senior Staff Performance Engineer, Firefox](https://hotfix.jobs/jobs/64810199-b6d8-436c-b310-10d723f9bffc) - Mozilla - Remote - CA$149k – CA$220k/yr
- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/b34c9349-da22-498a-9337-4a1f426a0b7f) - VGS - Remote - $145k – $260k/yr
- [Staff Engineer, Platform & Infrastructure](https://hotfix.jobs/jobs/48b953ba-a6fb-4833-9ad0-140dfe4943dc) - Nango - Remote - $140k – $220k/yr
- [Staff Platform Engineer](https://hotfix.jobs/jobs/37cd4cd2-d007-4013-8153-5443ae70f1dd) - Nango - Remote - $140k – $220k/yr
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/52f90907-0980-43e8-bd36-a8741c4039a6
**Canonical:** https://hotfix.jobs/jobs/52f90907-0980-43e8-bd36-a8741c4039a6