# Staff Software Engineer, Quality & Reliability

**Company:** [BuildOps](https://hotfix.jobs/companies/buildops)
**Location:** Los Angeles, CA, San Francisco, CA, Raleigh, NC, Toronto, Canada
**Role:** DevOps / SRE
**Salary:** CA$172k – CA$229k/yr
**Experience:** 7+ years
**Skills:** AWS, TypeScript, Java, Distributed Systems, Observability, Resilience Engineering, CI/CD, Release Safety, Automated Validation, Developer Platforms, Testability, Performance Engineering, Incident Learning
**Posted:** 2026-08-03

> Staff-level engineer who will establish company-wide strategy and platform capabilities for software quality, production reliability, observability, and safe delivery. The role requires cross-team technical leadership, distributed-systems experience, cloud expertise, and strong programming skills.

## Job Description

## Responsibilities
- Define and drive the technical strategy for engineering quality, production reliability, and safe software delivery.
- Identify systemic sources of customer-impacting failures and lead cross-team initiatives addressing root causes.
- Establish architectural principles, engineering standards, and paved roads for reliable system design and safe delivery.
- Partner with engineering teams on system and product design to improve resilience, operability, testability, and failure isolation.
- Build or guide shared platform capabilities for release safety, automated validation, production feedback, test data, environment management, and developer self-service.
- Advance observability to help teams understand system behavior, detect regressions, diagnose failures, and prioritize reliability investments.
- Improve validation across services, data boundaries, financial workflows, and other business-critical systems.
- Define meaningful quality and reliability measures and use them to demonstrate customer and engineering improvements.
- Lead technical programs spanning multiple teams and organizations, aligning stakeholders without direct authority.
- Mentor engineers and technical leaders on risk, reliability, and quality throughout the software lifecycle.
- Evaluate and evolve existing practices and technology.

## Success Measures
- Fewer customer-impacting defects and recurring production failures.
- Greater release confidence and lower change-failure rates.
- Faster failure detection, diagnosis, and recovery.
- Shorter, more reliable engineering feedback loops.
- Clearer ownership and visibility into critical system health.
- Increased engineering velocity without sacrificing safety or reliability.
- Broad adoption of shared practices without creating a centralized quality bottleneck.

## Requirements
- Significant software engineering experience, including Staff-level or equivalent scope on ambiguous, cross-cutting technical problems.
- Experience leading multi-team initiatives that improved production reliability, software delivery, platform capabilities, or engineering effectiveness.
- Strong systems thinking across architecture, data integrity, operations, developer workflows, and customer impact.
- Experience designing and operating distributed systems in a cloud environment such as AWS.
- Strong software design and programming skills in TypeScript, Java, or another relevant language.
- Experience with several of observability, resilience engineering, CI/CD, release safety, automated validation, developer platforms, testability, performance engineering, or incident learning.
- Ability to define useful engineering measures focused on outcomes.
- Success influencing architecture and engineering practices across teams without direct reporting relationships.
- Strong written and verbal communication.
- Practical approach balancing long-term direction with incremental improvements.

## Compensation & Benefits
- $172,000–$229,000 base salary, plus annual bonus and meaningful equity.
- Generous equity grant.
- Comprehensive benefits package.
- Flexible PTO and hybrid work schedules.
- One-time work-from-home allowance.
- Company events and team-building activities.
- Growth and career advancement opportunities.
- Hubs in Los Angeles, San Francisco, Toronto, and Raleigh, with hybrid schedules and lunch provided on in-office days.

## Similar jobs

- [Staff Site Reliability Engineer - Kubernetes](https://hotfix.jobs/jobs/903909fa-7a86-4f11-924e-bffc3234af22) - Okta - Bellevue, WA - $174k – $267k/yr
- [Staff Site Reliability Engineer, Kubernetes w/ active TS/SCI](https://hotfix.jobs/jobs/4b4d6848-5dfe-4852-9d84-7ec2d0a7f632) - Okta - Maryland - $174k – $238k/yr
- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/7f954a24-a52c-4bdc-82be-086a8926bf4d) - Okta - Bellevue, WA - $174k – $267k/yr
- [Staff Platform Engineer](https://hotfix.jobs/jobs/d1028610-5698-4ea7-850d-04183a6658da) - Motive - Buffalo, NY - $164k – $236k/yr
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/8c91b6be-3e6b-4298-8488-d5006f96a4f4
**Canonical:** https://hotfix.jobs/jobs/8c91b6be-3e6b-4298-8488-d5006f96a4f4