Skip to content
BuildOpsBuildOps

Staff Software Engineer, Quality & Reliability

Staff-level engineer who will establish company-wide strategy and platform capabilities for software quality, production reliability, observability, and safe delivery. The role requires cross-team technical leadership, distributed-systems experience, cloud expertise, and strong programming skills.

About the job

Responsibilities

  • Define and drive the technical strategy for engineering quality, production reliability, and safe software delivery.
  • Identify systemic sources of customer-impacting failures and lead cross-team initiatives addressing root causes.
  • Establish architectural principles, engineering standards, and paved roads for reliable system design and safe delivery.
  • Partner with engineering teams on system and product design to improve resilience, operability, testability, and failure isolation.
  • Build or guide shared platform capabilities for release safety, automated validation, production feedback, test data, environment management, and developer self-service.
  • Advance observability to help teams understand system behavior, detect regressions, diagnose failures, and prioritize reliability investments.
  • Improve validation across services, data boundaries, financial workflows, and other business-critical systems.
  • Define meaningful quality and reliability measures and use them to demonstrate customer and engineering improvements.
  • Lead technical programs spanning multiple teams and organizations, aligning stakeholders without direct authority.
  • Mentor engineers and technical leaders on risk, reliability, and quality throughout the software lifecycle.
  • Evaluate and evolve existing practices and technology.

Success Measures

  • Fewer customer-impacting defects and recurring production failures.
  • Greater release confidence and lower change-failure rates.
  • Faster failure detection, diagnosis, and recovery.
  • Shorter, more reliable engineering feedback loops.
  • Clearer ownership and visibility into critical system health.
  • Increased engineering velocity without sacrificing safety or reliability.
  • Broad adoption of shared practices without creating a centralized quality bottleneck.

Requirements

  • Significant software engineering experience, including Staff-level or equivalent scope on ambiguous, cross-cutting technical problems.
  • Experience leading multi-team initiatives that improved production reliability, software delivery, platform capabilities, or engineering effectiveness.
  • Strong systems thinking across architecture, data integrity, operations, developer workflows, and customer impact.
  • Experience designing and operating distributed systems in a cloud environment such as AWS.
  • Strong software design and programming skills in TypeScript, Java, or another relevant language.
  • Experience with several of observability, resilience engineering, CI/CD, release safety, automated validation, developer platforms, testability, performance engineering, or incident learning.
  • Ability to define useful engineering measures focused on outcomes.
  • Success influencing architecture and engineering practices across teams without direct reporting relationships.
  • Strong written and verbal communication.
  • Practical approach balancing long-term direction with incremental improvements.

Compensation & Benefits

  • $172,000–$229,000 base salary, plus annual bonus and meaningful equity.
  • Generous equity grant.
  • Comprehensive benefits package.
  • Flexible PTO and hybrid work schedules.
  • One-time work-from-home allowance.
  • Company events and team-building activities.
  • Growth and career advancement opportunities.
  • Hubs in Los Angeles, San Francisco, Toronto, and Raleigh, with hybrid schedules and lunch provided on in-office days.

Skills

AWS, TypeScript, Java, Distributed Systems, Observability, Resilience Engineering, CI/CD, Release Safety, Automated Validation, Developer Platforms, Testability, Performance Engineering, Incident Learning

Okta

Okta

Bellevue, WA
Staff Site Reliability Engineer - Kubernetes
$174k+/yrHybrid7+ YOEDevOps / SRE

Build and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.

Okta

Okta

Maryland
Staff Site Reliability Engineer, Kubernetes w/ active TS/SCI
$174k+/yrHybrid8+ YOEDevOps / SRE

Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.

Okta

Okta

Bellevue, WA
Staff Site Reliability Engineer
$174k+/yrHybrid7+ YOEDevOps / SRE

Leads reliability engineering for highly available, FedRAMP-compliant cloud services, including infrastructure architecture, automation, observability, incident response, and operational standards. Requires extensive Kubernetes, cloud, software engineering, and cross-team technical leadership experience, plus US-person eligibility and residence on US soil.

Motive

Motive

Buffalo, NY
Staff Platform Engineer
$164k+/yrOn-site7+ YOEDevOps / SRE

Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.