Skip to content
Shield AIShield AIUnited States

Sr. Staff Platform/Data Reliability Engineer, Databricks

Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.

180k – 270k/yr
Remote12+ YOEDevOps / SRE

About the role

Responsibilities

  • Own operational excellence for the Databricks platform, including monitoring, alerting, observability, incident response support, and production runbook patterns for data jobs and platform services.
  • Define and maintain CI/CD and promotion standards for Databricks assets, including workflows, jobs, notebooks, code packages, infrastructure configuration, and environment promotion from development to production.
  • Design and maintain platform standards for job orchestration, cluster and compute policies, service principal usage, environment isolation, and production execution reliability.
  • Establish reusable operational templates and enablement patterns for new domains onboarding to Databricks, including logging conventions, job tagging, metadata capture, and support handoff expectations.
  • Partner with the Senior Data Engineer to ensure ingestion and medallion patterns are observable, recoverable, cost-aware, and secure in production.
  • Work with the cloud and infrastructure team to align Databricks configuration and usage patterns with broader enterprise cloud standards, including commercial and future government-hosted environments.
  • Help enforce technical controls for data segregation, access boundaries, and operational compliance in a highly regulated environment.
  • Track and improve platform health metrics such as job success rates, incident trends, data pipeline reliability, cost efficiency, and environment drift.
  • Document platform standards, operational expectations, and support models so the Databricks platform can scale beyond a small founding team.
  • Mentor internal engineers developing platform responsibilities and Databricks operational expertise.

Requirements

  • 12+ years of relevant experience in data platform engineering, platform operations, site reliability engineering, or modern cloud data infrastructure.
  • Hands-on experience with Databricks or a closely related cloud data platform in production environments.
  • Experience designing or operating CI/CD, environment promotion, version control, and deployment automation for data platforms and pipelines.
  • Strong understanding of observability, monitoring, alerting, incident management, and reliability engineering.
  • Experience with compute policy design, workload isolation, service principals, and secure production execution patterns on cloud data platforms.
  • Ability to work effectively in regulated or security-sensitive environments with strong expectations around access control, auditability, and operational discipline.
  • Strong collaboration skills and comfort partnering with cloud/infrastructure, security, data engineering, and analytics stakeholders.

Nice-to-Haves

  • Databricks certification and/or demonstrated expertise with Delta Lake, Unity Catalog, Workflows, and Databricks Asset Bundles.
  • Experience with infrastructure as code and platform automation in enterprise environments.
  • Experience supporting commercial and government or otherwise segregated environments with different compliance and access requirements.
  • Experience in defense, aerospace, federal, or another regulated industry.

Skills

DatabricksDelta Lakeunity catalogdatabricks workflowsdatabricks asset bundlesCI/CDInfrastructure As CodeObservabilitysite reliability engineeringcloud data platformsversion controlcluster policiesservice principalsMonitoringIncident Management

Similar roles

DevOps / SRE jobs
Hightouch

Staff Engineer, AI Productivity

HightouchUnited States

Staff-level engineer building infrastructure, tooling, and documentation to make AI coding agents dramatically more productive across the codebase. Owns agentic dev environments, MCP integrations, and agent context.

180k – 400k/yrRemote7+ YOEDevOps / SRE
Gusto

Staff Software Engineer, AI Developer Tools

GustoDenver, CO +3

Staff-level engineer architecting AI-native developer tools and infrastructure to accelerate engineering velocity across Gusto. Requires 8+ years experience building production AI systems with deep expertise in LLMs, RAG, and multi-agent workflows.

180k – 245k/yrHybrid8+ YOEDevOps / SRE
Onebrief

Staff Infrastructure Engineer

OnebriefUnited States

Staff Infrastructure Engineer building and operating secure cloud-native and edge platforms for military collaboration software. Requires 5+ years production infrastructure experience, deep Kubernetes expertise, and ability to obtain SECRET clearance.

180k – 235k/yrRemote5+ YOEDevOps / SRE
Zocdoc

Staff Enterprise and Cloud Engineer

ZocdocNew York, NY

As a Staff Cloud IAM Engineer, you will own the technical vision and strategy for identity and access management across the corporate stack, focusing on Microsoft Entra ID, enterprise SSO/SCIM, and SaaS/AI platforms. You will design scalable identity governance, lead cross-functional initiatives, and ensure the reliability and security of core corporate infrastructure.

180k – 270k/yrRemote10+ YOEDevOps / SRE
xAI

Member Of Technical Staff - Cloud Infrastructure

xAIPalo Alto, CA +1

Designs, builds, and operates secure, scalable infrastructure including Kubernetes clusters and GPU hardware for large-scale AI workloads in classified US government environments. Requires 5+ years experience, Top Secret clearance, and expertise in IaC tools like Terraform and Ansible.

180k – 440k/yrOn-site5+ YOEDevOps / SRE