Staff Software Engineer, Quality & Reliability
Staff-level engineer who will establish company-wide strategy and platform capabilities for software quality, production reliability, observability, and safe delivery. The role requires cross-team technical leadership, distributed-systems experience, cloud expertise, and strong programming skills.
About the job
Responsibilities
- Define and drive the technical strategy for engineering quality, production reliability, and safe software delivery.
- Identify systemic sources of customer-impacting failures and lead cross-team initiatives addressing root causes.
- Establish architectural principles, engineering standards, and paved roads for reliable system design and safe delivery.
- Partner with engineering teams on system and product design to improve resilience, operability, testability, and failure isolation.
- Build or guide shared platform capabilities for release safety, automated validation, production feedback, test data, environment management, and developer self-service.
- Advance observability to help teams understand system behavior, detect regressions, diagnose failures, and prioritize reliability investments.
- Improve validation across services, data boundaries, financial workflows, and other business-critical systems.
- Define meaningful quality and reliability measures and use them to demonstrate customer and engineering improvements.
- Lead technical programs spanning multiple teams and organizations, aligning stakeholders without direct authority.
- Mentor engineers and technical leaders on risk, reliability, and quality throughout the software lifecycle.
- Evaluate and evolve existing practices and technology.
Success Measures
- Fewer customer-impacting defects and recurring production failures.
- Greater release confidence and lower change-failure rates.
- Faster failure detection, diagnosis, and recovery.
- Shorter, more reliable engineering feedback loops.
- Clearer ownership and visibility into critical system health.
- Increased engineering velocity without sacrificing safety or reliability.
- Broad adoption of shared practices without creating a centralized quality bottleneck.
Requirements
- Significant software engineering experience, including Staff-level or equivalent scope on ambiguous, cross-cutting technical problems.
- Experience leading multi-team initiatives that improved production reliability, software delivery, platform capabilities, or engineering effectiveness.
- Strong systems thinking across architecture, data integrity, operations, developer workflows, and customer impact.
- Experience designing and operating distributed systems in a cloud environment such as AWS.
- Strong software design and programming skills in TypeScript, Java, or another relevant language.
- Experience with several of observability, resilience engineering, CI/CD, release safety, automated validation, developer platforms, testability, performance engineering, or incident learning.
- Ability to define useful engineering measures focused on outcomes.
- Success influencing architecture and engineering practices across teams without direct reporting relationships.
- Strong written and verbal communication.
- Practical approach balancing long-term direction with incremental improvements.
Compensation & Benefits
- $172,000–$229,000 base salary, plus annual bonus and meaningful equity.
- Generous equity grant.
- Comprehensive benefits package.
- Flexible PTO and hybrid work schedules.
- One-time work-from-home allowance.
- Company events and team-building activities.
- Growth and career advancement opportunities.
- Hubs in Los Angeles, San Francisco, Toronto, and Raleigh, with hybrid schedules and lunch provided on in-office days.
Skills
AWS, TypeScript, Java, Distributed Systems, Observability, Resilience Engineering, CI/CD, Release Safety, Automated Validation, Developer Platforms, Testability, Performance Engineering, Incident Learning
Similar jobs
DevOps / SRE jobsBuild and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.
Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.
Leads reliability engineering for highly available, FedRAMP-compliant cloud services, including infrastructure architecture, automation, observability, incident response, and operational standards. Requires extensive Kubernetes, cloud, software engineering, and cross-team technical leadership experience, plus US-person eligibility and residence on US soil.
Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.