Staff Software Engineer I - SRE
Leads proactive reliability engineering and incident-management improvements for a large-scale, multi-cloud streaming platform. The role requires 10+ years in SRE, incident management, or reliability engineering, plus expertise in distributed systems, observability, Kubernetes, cloud infrastructure, and incident tooling.
About the job
Responsibilities
Proactive Reliability Engineering
- Analyze systemic failure patterns and design improvements that prevent incident recurrence.
- Define and maintain SLO/SLA frameworks and use error budgets to guide reliability investments.
- Build tooling and automation to reduce incident-response toil and scale team impact.
- Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack.
- Analyze reliability data to identify systemic improvements and build dashboards that drive action.
- Explore AI-assisted approaches to documentation quality and incident analysis.
- Design scalable reliability standards that reduce reactive workload over time.
Incident Management Program
- Own standards, practices, and continuous improvement of incident response.
- Serve as an on-call Incident Commander for production incidents, including as escalation Incident Commander when incidents exceed a team’s management chain.
- Develop and deliver training programs for engineering teams at all levels.
- Coach teams through post-mortems and the development of actionable corrective actions.
Customer Root Cause Analysis
- Edit and review customer-facing incident documents for quality and clarity.
- Drive turnaround SLAs while maintaining technical accuracy.
- Ensure clear explanations of what happened, why it happened, and how recurrence will be prevented.
Cross-Team Leadership
- Partner with engineering leaders to elevate reliability practices.
- Serve as an expert whom teams proactively engage for guidance.
Requirements
- 10+ years of experience in SRE, incident management, or reliability engineering.
- Cloud experience with at least one of AWS, Google Cloud, or Azure.
- Deep expertise with incident-management tooling such as Rootly, PagerDuty, or similar platforms.
- Strong understanding of distributed systems and failure modes at scale; Kafka or event-streaming expertise is preferred, or demonstrated ability to quickly master complex systems.
- Deep experience with observability, including metrics, logging, and tracing, with the ability to diagnose complex issues.
- Kubernetes and container-orchestration experience.
- Understanding of CI/CD pipelines and release processes.
- Systems thinking and understanding of how infrastructure design choices affect failure modes and recovery.
- Familiarity with SLO/SLA frameworks.
- Track record as a trusted advisor across engineering organizations.
- Experience driving organization-wide process and cultural changes.
- Strong written communication skills for design documents, one-pagers, and runbooks.
- Post-mortem facilitation experience.
- Experience with asynchronous collaboration across time zones.
- Experience navigating reliability and incident programs in organizations with 500+ engineers.
Nice to Have
- Multi-cloud experience with at least two of AWS, Google Cloud, or Azure.
- Experience with modern CI/CD, GitHub, and AI-assisted workflows.
Compensation and Benefits
- Global follow-the-sun coverage with clean handoffs designed to support sustainable working hours.
Skills
SRE, Incident Management, AWS, GCP, Azure, Rootly, Pagerduty, Distributed Systems, Kafka, Observability, Kubernetes, CI/CD, Slo/Sla, Error Budgets, Jira
Similar jobs
DevOps / SRE jobsOwn reproducible C++ build and release workflows, containerized CI/CD, artifact promotion, and edge deployment for autonomy and vision products. The role requires strong production C++, build-system, container, Linux, and constrained-environment delivery experience.
Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.
Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.