Incident Manager
Lead high-severity production incidents as commander, drive root cause analysis and reliability fixes across cloud systems, and deliver clear technical and customer communications. Requires 5+ years in SRE/incident ops, cloud and observability expertise, scripting skills, and strong writing ability.
About the job
Impact
- Lead critical incidents — coordinate multi-disciplinary response efforts across Databricks’ cloud-based services to rapidly mitigate impact and restore operations.
- Drive technical root cause analysis and reliability improvements: collaborate with engineering teams to trace and document underlying causes across distributed systems, services, and data stores. Summarize key learnings, clearly communicate action items, and ensure that technical and procedural improvements are followed through.
- Own communications during incidents — deliver frequent, high-quality updates to internal stakeholders (executives, engineering leadership, support) and compose and publish customer-facing notifications that are accurate, timely, and empathetic.
- Mentor and train peers in both incident communication and technical response disciplines to raise the overall quality of Databricks’ incident response.
Requirements
- 5+ years of experience in incident management, site reliability engineering, or production operations supporting large-scale, cloud-native systems.
- Proven ability to lead and coordinate high-severity incidents, including identifying impact, isolating fault domains, and managing multi-team response efforts.
- Strong understanding of cloud infrastructure (AWS, Azure, or GCP) — including compute, networking, storage, and observability components.
- Deep expertise in log analysis and debugging: familiarity with log aggregation and search tools (e.g., Datadog, Elasticsearch, Splunk, Cloud Logging, or OpenTelemetry).
- Hands-on experience with observability systems — metrics, logging, and tracing frameworks (Prometheus, Grafana, OpenTelemetry, etc.).
- Proficiency in at least one major programming or scripting language (Python, Go, or Bash) for automating diagnostics, data collection, or analysis.
- Experience developing and maintaining incident playbooks and communication templates to ensure consistent, timely updates.
- Excellent contextual interpretation and writing skills, as well as the ability to effectively summarize and communicate to both technical and business audiences.
- BS, Master's or other advanced degree in Computer Science or Computer Engineering, or related Engineering field.
Skills
Incident Management, Site Reliability Engineering, AWS, Azure, GCP, Datadog, Elasticsearch, Splunk, Prometheus, Grafana, OpenTelemetry, Python, Go, Bash
Similar jobs
Support Engineering jobsInvestigates and resolves high-impact financial institution integration issues by developing, deploying, and monitoring code while coordinating with external data partners. Requires 1.5+ years in a technical role, API or integration experience, strong debugging, SQL-based analysis, and clear communication.
Provides advanced technical, operational, and field support for Skydio UAS hardware, docks, cloud systems, and networking. The role requires at least three years of UAS flight experience, strong troubleshooting skills, and regional travel of up to 30–50%.
Leads technical support operations by managing support engineers, architecting integrations and automation, and owning high-severity escalations and operational reliability. Requires 5+ years managing technical support teams, strong systems expertise, and hands-on experience shipping automation and AI workflows.
The Support Engineer troubleshoots GitLab deployments for U.S. government and public-sector customers in secure, air-gapped, and regulated environments. The role combines Linux systems investigation, scripting, GitLab internals, customer support, documentation, automation, and cross-functional engineering collaboration.
This role solves complex technical customer problems by debugging APIs, databases, logs, code, and production systems while partnering with Product and Engineering. It requires strong Python and SQL skills, clear communication, customer empathy, and a builder mindset; hardware-development experience is a plus.