Sr. Staff Technical Program Manager - Reliability
Leads reliability strategy and execution for Databricks' multi-cloud infrastructure, partnering with senior engineering leaders to drive roadmaps, programs, and best practices. Requires 10+ years in cloud/SRE with hyperscale experience.
About the job
Responsibilities
Lead Reliability Strategy + Multi-Quarter Roadmaps
- Partner with senior engineering leadership to define long-term Reliability roadmap and ensure alignment across Platform Engineering, Compute Fleet Management, SRE, Security, and Cloud Partnerships.
Drive Execution of Critical Reliability Programs
- Own end-to-end program execution: planning, risk management, dependency mapping, trade-off decisions, status reporting, and delivery.
- Identify process/architecture gaps and drive improvements with Tech Leads.
Partner Deeply with Engineering & Influence Technical Direction
- Leverage infrastructure/SRE background to guide design and prioritization.
- Facilitate cross-functional alignment and apply systems thinking to improve scalability, fault tolerance, automation, and tooling.
Elevate Reliability Culture
- Drive adoption of best practices: error budgets, incident reviews, design-for-resilience, operational readiness.
- Implement governance, processes, metrics, and documentation.
Required Experience & Qualifications
- 10+ years managing large-scale technical programs in cloud infrastructure, distributed systems, SRE, or platform engineering.
- Experience with 2+ hyperscale clouds (AWS, Azure, GCP), multi-AZ/region architecture.
- Success leading Reliability Programs (availability, failover, incident reduction).
- Strong understanding of infrastructure/distributed systems/SRE; engineering/SRE experience preferred.
- Partnering with senior leadership on strategy and multi-team initiatives.
- Translate ambiguous goals into plans with milestones/KPIs.
- Manage cross-org dependencies, risks, multi-quarter timelines.
- Delivering programs across multiple clouds/cloud-native services.
- Building/scaling engineering processes and frameworks.
Preferred Qualifications
- Background in distributed systems engineering, SRE, platform infrastructure, or cloud services.
- Experience with compute fleets, container orchestration, autoscaling, control-plane.
- Familiarity with SLOs, error budgets, chaos engineering, incident management.
- Expertise with Jira or equivalent.
- Bachelor’s in CS/Engineering or related; advanced degree preferred.
Skills
AWS, Azure, GCP, Distributed Systems, SRE, Jira, SLOs, Error Budgets, Chaos Engineering, Container Orchestration
Similar jobs
Technical Program Management jobsLeads high-impact technical programs for the Duolingo English Test, coordinating engineering, product, data, legal, security, and external stakeholders from planning through launch. Requires staff-level autonomy, strong technical communication, and experience managing complex programs in agile software environments.
Owns the strategy, governance, architecture, and hands-on development of Fetch’s People data and analytics foundation. The role builds semantic models, reporting products, and operating processes while enabling trusted metrics, self-service analytics, and future AI automation.
Owns the electrical and electrical-distribution program scope for automotive vehicle programs developed with external partners, coordinating harness, E/E integration, validation, release control, milestones, and executive reporting. Requires 10+ years in automotive technical program management or related engineering leadership.
Leads company-scale search signals programs spanning offsite data, content understanding, relevance, and AI adoption. The role requires 10+ years of technical or analytical experience, strong cross-functional influence, and expertise in search, recommendation systems, ML/AI, data integration, and AI governance.
Leads cross-functional programs delivering end-to-end autonomous mobility features across vehicle, autonomy, and operations teams. Requires 10+ years in engineering or program management, experience with complex product development, and strong stakeholder communication.