Skip to content
OktaOkta

Staff Software Reliability Engineer - Data Platform

The Staff Software Reliability Engineer will design, build, optimize, and operate scalable streaming and distributed data-platform infrastructure supporting analytics and machine learning. The role requires 5+ years of industry experience, strong software engineering skills, and expertise with distributed data technologies and reliability practices.

About the job

Responsibilities

  • Design, implement, and own data-intensive, high-performance, scalable platform components.
  • Collaborate with engineering teams, architects, and cross-functional partners on project development, design, and implementation.
  • Conduct and participate in design reviews, code reviews, analysis, and performance tuning.
  • Coach and mentor engineers.
  • Debug production issues across services and multiple levels of the stack.
  • Participate in the on-call rotation and incident management.

Requirements

  • 5+ years of industry experience.
  • 2+ years of experience with an object-oriented language, preferably Java.
  • Hands-on experience with cloud-based distributed computing technologies, including:
    • Messaging systems such as Kinesis and Kafka.
    • Data processing systems such as Flink, Spark, and Beam.
    • Storage and compute systems such as Snowflake, Databricks, and Hadoop.
    • Coordinators and schedulers such as those in Kubernetes, Hadoop, and Mesos.
  • Experience developing and tuning highly scalable distributed systems.
  • Excellent grasp of software engineering principles.
  • Solid understanding of multithreading, garbage collection, and memory management.
  • Experience with reliability engineering, particularly data quality, data observability, and incident management.

Nice to Have

  • Experience maintaining security, encryption, identity management, or authentication infrastructure.
  • Experience using major public cloud providers to build mission-critical, high-volume services.
  • Experience developing data integration applications for large-scale, petabyte-scale environments across batch and online systems.
  • Contributions to distributed systems or experience using high-volume or critical systems such as Kafka or Hadoop.
  • Experience developing Kubernetes-based services on AWS.

Compensation and Benefits

  • Annual base salary for candidates located in Canada: $160,000–$220,000 CAD.
  • Equity, where applicable, bonus, and benefits including health, dental, and vision insurance, RRSP matching, healthcare spending, telemedicine, and paid leave including PTO and parental leave.

Skills

Java, Kinesis, Kafka, Flink, Spark, Apache Beam, Snowflake, Databricks, Hadoop, Kubernetes, Mesos, AWS, Distributed Systems, Multithreading, Incident Management

Mozilla

Mozilla

Canada

Senior Staff Performance Engineer, Firefox
CA$149k+/yrRemote7+ YOEDevOps / SRE

Leads Firefox performance engineering by writing code, profiling bottlenecks, improving benchmarks, and guiding cross-functional teams. Requires 7+ years of experience, strong C++ and JavaScript skills, and expertise in performance-critical software, profiling, concurrency, and systems analysis.

VGS

VGS

United States
Staff Infrastructure Engineer
$145k+/yrRemote8+ YOEDevOps / SRE

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting mission-critical payment systems. Requires 8+ years of distributed-systems experience and deep expertise in infrastructure as code, Kubernetes, automation, and cloud networking.

Nango

Nango

United States
Staff Engineer, Platform & Infrastructure
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.

Nango

Nango

United States
Staff Platform Engineer
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.