# Senior Platform Monitoring Engineer

**Company:** [Databricks](https://hotfix.jobs/companies/databricks)
**Location:** Unspecified
**Role:** DevOps / SRE
**Experience:** 6+ years
**Skills:** AWS, Azure, GCP, Docker, Kubernetes, Elk, Prometheus, Grafana, Pagerduty, Python, Observability, Incident Response
**Posted:** 2026-08-27

> Leads platform observability, monitoring, and incident response efforts, designing alerting and automation that improve reliability and customer experience. Requires 6+ years in SRE, DevOps, production engineering, or a similar role, along with cloud, container orchestration, and monitoring expertise.

## Job Description

## Responsibilities
- Lead platform incident investigations, coordinating cross-functional teams through detection, mitigation, and resolution to minimize customer impact.
- Conduct post-incident root cause analyses across infrastructure, services, and cloud providers; identify systemic patterns and prevention measures.
- Design and implement customer-focused alerting pipelines and end-to-end observability workflows.
- Build automation tools, establish reusable monitoring patterns, and address reliability gaps affecting customer experience.
- Mentor junior engineers on observability patterns, alert design, and service health metrics.
- Participate in an on-call rotation.

## Requirements
- At least 6 years of experience as an SRE, DevOps Engineer, Production Engineer, or similar.
- Production experience with at least one major cloud provider: AWS, Azure, or Google Cloud.
- Proficiency with Docker and Kubernetes.
- Hands-on experience with monitoring, logging, and alerting tools such as ELK, Prometheus, Grafana, and PagerDuty.
- Ability to architect monitoring solutions that correlate metrics, logs, and traces.
- Strong proficiency in Python or a similar programming language, with the ability to build production-quality automation tools.
- Experience owning incident lifecycles from detection through resolution and post-mortem analysis in demanding production environments.
- Bachelor's, master's, or doctoral degree in Computer Science, Computer Engineering, or a related engineering field.

## Similar jobs

- [Senior Platform Engineer](https://hotfix.jobs/jobs/71245322-afd8-4019-8525-65088fed1493) - Shield AI - San Diego, CA - $141k – $212k/yr
- [Senior Network Engineer](https://hotfix.jobs/jobs/d4ecdaa5-c0ab-49f3-baed-0b13deaaf6c0) - Shield AI - San Mateo, CA - $140k – $211k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/3317571d-7a67-4059-b23e-af1a9033cf96) - Astra - Remote - $190k – $230k/yr
- [Senior Software Engineer, Core Infra Systems](https://hotfix.jobs/jobs/6261dc33-aecd-4734-b00d-46f60f4f38c9) - Coinbase - Remote - $186k – $219k/yr
- [Senior Site Infrastructure Engineer](https://hotfix.jobs/jobs/847213bd-91b8-44f9-b450-477e4aa3714e) - Shield AI - Seattle, WA - $110k – $210k/yr

**Apply:** https://hotfix.jobs/jobs/74e0499a-c2ba-4018-8af8-1f4814399ebb
**Canonical:** https://hotfix.jobs/jobs/74e0499a-c2ba-4018-8af8-1f4814399ebb