Engineering Manager - Compute Infra
Leads engineering team building and operating large-scale compute infrastructure platform supporting Databricks workloads across clouds. Requires 3+ years management and 10+ years in distributed systems with cloud/container expertise.
About the job
Responsibilities
- Own and evolve the compute platform to support all Databricks workloads, enabling high-velocity product development and best-in-class performance.
- Hire exceptional engineers and invest in their growth through coaching, feedback, and career development.
- Raise the technical and operational bar through strong design practices, testing, and a culture of engineering excellence and platform mindset.
- Partner with engineering and product leadership to define long-term strategy and roadmaps.
- Lead cross-team initiatives spanning product and infrastructure.
- Influence architectural decisions beyond your immediate team.
Requirements
- 10+ years of experience building and operating large-scale distributed systems, including strong practices around testing, monitoring, and reliability.
- 3+ years of engineering management experience leading high-performing teams.
- Deep expertise in cloud infrastructure (AWS, Azure, GCP) and containerized environments (Kubernetes) preferred.
- Proven ability to collaborate across engineering, product, and business stakeholders.
- Experience scaling teams and systems through rapid growth.
- BS, MS, or PhD in Computer Science or related field, or equivalent experience.
Skills
Kubernetes, AWS, Azure, GCP, Distributed Systems
Similar jobs
Engineering Management jobsLeads the engineering team responsible for property-management platform capabilities, data models, integrations, APIs, and production reliability. Requires 5+ years of software engineering experience, technical leadership, backend expertise, and strong database and distributed-systems knowledge.
Leads the engineering team responsible for launching and scaling Credential Governance, a product that helps enterprises discover, control, and manage company-owned credentials. Requires several years of engineering management experience, production-service leadership, technical fluency, and experience growing remote teams.
Leads the Coordination Systems Storage team, managing five engineers while contributing to distributed-systems development, incident response, and reliability initiatives. Requires platform infrastructure experience, strong coding and performance skills, and a background in operational excellence.
Leads the Metadata engineering team responsible for discovery APIs, catalog, lineage, and run-history services. The manager owns delivery, architecture, reliability, operational health, stakeholder planning, and engineer development in a distributed cloud environment.
Leads a technical engineering team building automated IT infrastructure, inventory systems, and device-lifecycle software. The role requires engineering management experience, strong distributed-systems and architecture expertise, and hands-on technical leadership in an ambiguous environment.