# Staff Software Engineer I - SRE

**Company:** [Confluent](https://hotfix.jobs/companies/confluent)
**Location:** Unspecified
**Role:** DevOps / SRE
**Experience:** 10+ years
**Skills:** SRE, Incident Management, AWS, GCP, Azure, Rootly, Pagerduty, Distributed Systems, Kafka, Observability, Kubernetes, CI/CD, Slo/Sla, Error Budgets, Jira
**Posted:** 2026-02-16

> Leads proactive reliability engineering and incident-management improvements for a large-scale, multi-cloud streaming platform. The role requires 10+ years in SRE, incident management, or reliability engineering, plus expertise in distributed systems, observability, Kubernetes, cloud infrastructure, and incident tooling.

## Job Description

## Responsibilities

### Proactive Reliability Engineering
- Analyze systemic failure patterns and design improvements that prevent incident recurrence.
- Define and maintain SLO/SLA frameworks and use error budgets to guide reliability investments.
- Build tooling and automation to reduce incident-response toil and scale team impact.
- Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack.
- Analyze reliability data to identify systemic improvements and build dashboards that drive action.
- Explore AI-assisted approaches to documentation quality and incident analysis.
- Design scalable reliability standards that reduce reactive workload over time.

### Incident Management Program
- Own standards, practices, and continuous improvement of incident response.
- Serve as an on-call Incident Commander for production incidents, including as escalation Incident Commander when incidents exceed a team’s management chain.
- Develop and deliver training programs for engineering teams at all levels.
- Coach teams through post-mortems and the development of actionable corrective actions.

### Customer Root Cause Analysis
- Edit and review customer-facing incident documents for quality and clarity.
- Drive turnaround SLAs while maintaining technical accuracy.
- Ensure clear explanations of what happened, why it happened, and how recurrence will be prevented.

### Cross-Team Leadership
- Partner with engineering leaders to elevate reliability practices.
- Serve as an expert whom teams proactively engage for guidance.

## Requirements
- 10+ years of experience in SRE, incident management, or reliability engineering.
- Cloud experience with at least one of AWS, Google Cloud, or Azure.
- Deep expertise with incident-management tooling such as Rootly, PagerDuty, or similar platforms.
- Strong understanding of distributed systems and failure modes at scale; Kafka or event-streaming expertise is preferred, or demonstrated ability to quickly master complex systems.
- Deep experience with observability, including metrics, logging, and tracing, with the ability to diagnose complex issues.
- Kubernetes and container-orchestration experience.
- Understanding of CI/CD pipelines and release processes.
- Systems thinking and understanding of how infrastructure design choices affect failure modes and recovery.
- Familiarity with SLO/SLA frameworks.
- Track record as a trusted advisor across engineering organizations.
- Experience driving organization-wide process and cultural changes.
- Strong written communication skills for design documents, one-pagers, and runbooks.
- Post-mortem facilitation experience.
- Experience with asynchronous collaboration across time zones.
- Experience navigating reliability and incident programs in organizations with 500+ engineers.

## Nice to Have
- Multi-cloud experience with at least two of AWS, Google Cloud, or Azure.
- Experience with modern CI/CD, GitHub, and AI-assisted workflows.

## Compensation and Benefits
- Global follow-the-sun coverage with clean handoffs designed to support sustainable working hours.

## Similar jobs

- [Staff C++ Build and Release Engineer](https://hotfix.jobs/jobs/1c172edb-d2a7-49a8-a5c0-7442664e4c8b) - Shield AI - Melbourne, Australia
- [Staff+ Site Reliability Engineer, Safeguards ML Infra](https://hotfix.jobs/jobs/6492550a-2ff4-498b-8247-470adae7d0c3) - Anthropic - San Francisco, CA - $320k – $485k/yr
- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/6e086905-4170-407a-a7a4-2e05df0701c5) - Polymarket - New York, NY - $250k – $500k/yr
- [Senior/Staff Infrastructure & Platform Engineer](https://hotfix.jobs/jobs/503b2675-718e-49a4-9692-0ca04a50b707) - Fortanix - Santa Clara, CA - $155k – $230k/yr
- [Staff Network Engineer, App Platform](https://hotfix.jobs/jobs/a410525c-f62d-4037-a6a5-40fa01a905e7) - Scale AI - San Francisco, CA

**Apply:** https://hotfix.jobs/jobs/7cc3d15c-cd32-4373-a275-47bc71130cf7
**Canonical:** https://hotfix.jobs/jobs/7cc3d15c-cd32-4373-a275-47bc71130cf7