Staff Site Reliability Engineer, Release Engineering
Staff SRE on the Release Engineering team defining and scaling reliability practices, architecting SLO/error-budget programs, and driving progressive delivery and automated safety gates across product engineering.
About the job
What excites you
- Lead the expansion of reliability standards across product engineering, converting foundational infrastructure into lasting operational habits and tooling.
- Architect and manage the SLO and error-budget framework, empowering teams to utilize reliability data for strategic product and release choices.
- Promote widespread use of progressive delivery and automated safety gates, ensuring high velocity without compromising production stability.
- Guide emerging product teams toward production readiness through expertise in observability, incident response, and scalable deployment health.
- Collaborate with SRE, Platform, and Infrastructure teams to transform complex production requirements into intuitive, self-service platform features.
- Direct the response to critical incidents and ensure the resulting post-mortem actions yield permanent improvements to the platform.
- Prepare for an AI-driven development landscape by scaling our safety nets to handle an increased volume and frequency of code changes.
What excites us
- Over 8 years of professional experience in backend systems, SRE, or platform engineering roles.
- Proven track record of designing reliability programs—such as service maturity models or SLI frameworks—that achieved cross-team adoption.
- Direct experience building or operating canary rollout systems, metric-gated analysis, or automated rollback infrastructure.
- Technical proficiency in software development, with a preference for Go or similar systems languages.
- Ability to drive organizational change and influence engineering culture without formal authority.
- Sound technical judgment in high-stakes production scenarios, balancing user impact with developer velocity.
- Prior exposure to Kubernetes, service mesh technologies, Prometheus, or ArgoCD is considered a strong asset.
Compensation and Benefits
- Additional compensation in the form(s) of equity and/or commission are dependent on the position offered.
- Plaid provides a comprehensive benefit plan, including medical, dental, vision, and 401(k).
Skills
Go, Kubernetes, Prometheus, Argo CD, Service Mesh, Slo, Error Budgets, Canary Rollouts, Observability, Incident Response
Similar jobs
DevOps / SRE jobsLeads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.
Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.
Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.