Staff Software Engineer
Leads the design and operation of Databricks’ large-scale Data Intelligence Platform, including metrics stores, ETL frameworks, multi-cloud pipelines, governance, and infrastructure tooling. Requires extensive industry experience, distributed-systems expertise, and technical leadership across complex data infrastructure initiatives.
About the job
Responsibilities
- Design and run the Databricks metrics store for sharing and aggregating detailed metrics across business units and engineering teams, with high quality, introspection, and query performance.
- Design and run the cross-company Data Intelligence Platform containing business and product metrics.
- Develop tooling and infrastructure to manage and run Databricks on Databricks at scale across multiple clouds, geographies, and deployment types.
- Build CI/CD processes, pipeline test frameworks, data-quality tooling, and infrastructure-as-code tooling.
- Design the base ETL framework used by company data pipelines.
- Partner with engineering teams to develop the long-term vision and requirements for Databricks products.
- Build reliable data pipelines and solve data problems using Databricks, partner products, and open-source tools.
- Establish conventions and APIs for telemetry, debugging, feature, and audit-event log data.
- Represent Databricks at academic and industry conferences and events.
Requirements
- 12+ years of industry experience.
- 4+ years of experience building large-scale distributed systems.
- 5+ years providing technical leadership on large projects involving ETL frameworks, metrics stores, infrastructure management, or data security.
- Experience building, shipping, and operating reliable multi-geography data pipelines at scale.
- Experience operating workflow or orchestration frameworks, including Airflow, dbt, or commercial enterprise tools.
- Experience with large-scale messaging systems such as Kafka, RabbitMQ, or commercial systems.
- Excellent cross-functional and communication skills, with the ability to build consensus.
- Passion for data infrastructure and enabling others to access data more easily.
Skills
Databricks, Data Intelligence Platform, Metrics Store, ETL, Data Pipelines, Distributed Systems, Airflow, dbt, Kafka, RabbitMQ, CI/CD, Infrastructure As Code, Data Quality, Data Governance, Multi-Cloud
Similar jobs
Data Engineering jobsLeads the design and development of scalable analytic data infrastructure, including the foundation for Celonis’s Digital Twin. Requires 10+ years of analytics or data engineering experience, enterprise Databricks expertise, and strong stakeholder communication.
Leads the re-platforming of Vanta’s compliance data layer from MongoDB to schema-aware PostgreSQL across high-throughput Kafka and S3 pipelines. The role requires staff-level distributed systems expertise, migration leadership, and strong experience with relational and document data modeling.
Leads the design, operation, and evolution of Databricks’ cross-company Data Intelligence Platform, including large-scale data systems, pipelines, governance, and infrastructure. Requires 10+ years of distributed-systems experience and substantial technical leadership on production data platforms.
Build and operate scalable lakehouse infrastructure, streaming and CDC pipelines, query systems, and self-serve BI capabilities. Requires 5+ years of data engineering experience, strong Kubernetes and infrastructure-as-code expertise, and hands-on experience with distributed data platforms.
Build and operate distributed systems powering Apache Pinot’s real-time analytics platform at massive scale. The role requires strong distributed-systems expertise, end-to-end delivery ownership, and a focus on reliability, observability, and performance.