Director of Engineering, Organizations & Cells
Leads the engineering organization responsible for scaling GitLab.com through cell-based architecture, customer migrations, routing, and multi-cloud infrastructure. The role requires director-level people leadership, distributed-systems expertise, large-scale database and migration experience, and strong asynchronous communication.
About the job
Responsibilities
- Deliver the Cells and Organizations roadmap, including cell provisioning and lifecycle, routing and topology, organization data migration tooling, and platform feature parity.
- Execute cohort-based customer Organization migrations onto cells with zero data loss, verified rollback, and no downtime outside the tenant in flight.
- Shape a multi-cloud, multi-region cell footprint and decide what remains cell-local versus what crosses boundaries.
- Build and lead the engineering organization by hiring, leveling, and developing managers and senior individual contributors across several teams.
- Set the technical bar and operating cadence.
- Drive commitments and sequencing with peer engineering leaders, Product, Infrastructure, and Security across time zones and asynchronously.
Success Measures
- Customer Organizations running on cells in production, migrated cleanly and on schedule.
- Availability and latency improvements from isolation and load distribution.
- Infrastructure cost per unit of capacity declining as the footprint grows.
- Throughput and retention of the teams built.
Requirements
- Experience leading multiple teams through managers, roughly 25–60 engineers, including hiring and developing managers.
- Track record of shipping multi-quarter infrastructure programs with hard external dependencies.
- Ability to operate as a peer to VPs and staff-plus engineers, make architectural decisions on technical merit, and adapt based on evidence.
- Strong written communication for a remote, asynchronous work environment.
- Experience building and operating large-scale distributed systems in production, such as multi-tenant SaaS, sharded or cell-based architectures, control planes, or data platforms at comparable scale.
- Working fluency in relational database scaling, including partitioning, sharding, replication, large-scale data migration, and request routing.
- Operational experience with incident command, capacity planning, and blast-radius containment.
- Deep cloud infrastructure experience with at least one major provider, ideally more than one.
- Understanding of how agentic access patterns affect platform read volume, concurrency, and authorization models.
- Experience using AI tooling in engineering practice and raising its adoption across a team.
Benefits
- Flexible paid time off
- Team member resource groups
- Equity compensation and employee stock purchase plan
- Growth and development fund
- Parental leave
Skills
Distributed Systems, Multi-Tenant Saas, Database Sharding, Database Replication, Data Migration, Request Routing, Multi-Cloud Infrastructure, Multi-Region Architecture, Incident Command, Capacity Planning, AI Tools
Similar jobs
Engineering Management jobsLeads and builds Starburst’s AI agent platform in India, owning its roadmap, architecture, production operations, and evolution from internal infrastructure to enterprise-facing products. Requires 10+ years of engineering experience, team leadership, platform development, and production LLM systems expertise.
Leads India-based Site Reliability Engineering teams responsible for Okta’s cloud platform, databases, networking, Kubernetes, CI/CD, observability, FinOps, and automation. Requires 16+ years in infrastructure or SRE, substantial people-management experience, and expertise in AWS, Kubernetes, Terraform, and reliable SaaS operations.
Leads a team of Partner Solution Architects across the APAC partner ecosystem, driving strategic relationships, technical alignment, joint solutions, and business growth. Requires extensive partner-management or solution-architecture experience, substantial leadership experience, and a bachelor’s degree or equivalent.
Leads the organization and strategy for Databricks’ large-scale data infrastructure, including billing correctness, disaster recovery, reliability, deployment automation, and data integration. The role requires extensive distributed-systems experience, infrastructure leadership, and experience managing managers.
Leads engineering organization and developer-platform initiatives, including scalable cloud services, AI-assisted development, technical strategy, hiring, and leadership development. Requires 15+ years building distributed systems and experience managing high-performance engineering teams.