# Sr. Staff Platform/Data Reliability Engineer, Databricks

**Company:** [Shield AI](https://hotfix.jobs/companies/shield-ai)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $180k – $270k/yr
**Experience:** 12+ years
**Skills:** Databricks, Delta Lake, unity catalog, databricks workflows, databricks asset bundles, CI/CD, Infrastructure As Code, Observability, site reliability engineering, cloud data platforms, version control, cluster policies, service principals, Monitoring, Incident Management
**Posted:** 2026-08-12

> Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.

## Job Description

## Responsibilities

- Own operational excellence for the Databricks platform, including monitoring, alerting, observability, incident response support, and production runbook patterns for data jobs and platform services.
- Define and maintain CI/CD and promotion standards for Databricks assets, including workflows, jobs, notebooks, code packages, infrastructure configuration, and environment promotion from development to production.
- Design and maintain platform standards for job orchestration, cluster and compute policies, service principal usage, environment isolation, and production execution reliability.
- Establish reusable operational templates and enablement patterns for new domains onboarding to Databricks, including logging conventions, job tagging, metadata capture, and support handoff expectations.
- Partner with the Senior Data Engineer to ensure ingestion and medallion patterns are observable, recoverable, cost-aware, and secure in production.
- Work with the cloud and infrastructure team to align Databricks configuration and usage patterns with broader enterprise cloud standards, including commercial and future government-hosted environments.
- Help enforce technical controls for data segregation, access boundaries, and operational compliance in a highly regulated environment.
- Track and improve platform health metrics such as job success rates, incident trends, data pipeline reliability, cost efficiency, and environment drift.
- Document platform standards, operational expectations, and support models so the Databricks platform can scale beyond a small founding team.
- Mentor internal engineers developing platform responsibilities and Databricks operational expertise.

## Requirements

- 12+ years of relevant experience in data platform engineering, platform operations, site reliability engineering, or modern cloud data infrastructure.
- Hands-on experience with Databricks or a closely related cloud data platform in production environments.
- Experience designing or operating CI/CD, environment promotion, version control, and deployment automation for data platforms and pipelines.
- Strong understanding of observability, monitoring, alerting, incident management, and reliability engineering.
- Experience with compute policy design, workload isolation, service principals, and secure production execution patterns on cloud data platforms.
- Ability to work effectively in regulated or security-sensitive environments with strong expectations around access control, auditability, and operational discipline.
- Strong collaboration skills and comfort partnering with cloud/infrastructure, security, data engineering, and analytics stakeholders.

## Nice-to-Haves

- Databricks certification and/or demonstrated expertise with Delta Lake, Unity Catalog, Workflows, and Databricks Asset Bundles.
- Experience with infrastructure as code and platform automation in enterprise environments.
- Experience supporting commercial and government or otherwise segregated environments with different compliance and access requirements.
- Experience in defense, aerospace, federal, or another regulated industry.

## Similar roles

- [Staff Engineer, AI Productivity](https://hotfix.jobs/jobs/8dcfdcba-4e0d-400e-a95e-d01acbf151e3) - Hightouch - Remote - $180k – $400k/yr
- [Staff Software Engineer, AI Developer Tools](https://hotfix.jobs/jobs/4ad0f960-2ad7-4648-829d-c4aff746157f) - Gusto - Denver, CO - $180k – $245k/yr
- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/31b641c0-5800-45dd-bdac-4948a91afb52) - Onebrief - Remote - $180k – $235k/yr
- [Staff Enterprise and Cloud Engineer](https://hotfix.jobs/jobs/7cb251e8-eccf-45c1-b5b4-12896c8df1e0) - Zocdoc - Remote - $180k – $270k/yr
- [Member Of Technical Staff - Cloud Infrastructure](https://hotfix.jobs/jobs/66bad3a4-88a7-434f-a081-9be91e2db188) - xAI - Palo Alto, CA - $180k – $440k/yr

**Apply:** https://hotfix.jobs/jobs/b1ca8d0a-2503-4891-8213-a0c0ca8c01ad
**Canonical:** https://hotfix.jobs/jobs/b1ca8d0a-2503-4891-8213-a0c0ca8c01ad