Director, Software Engineering
Lead ClickUp's Cloud Platform team to build scalable, cost-efficient AWS infrastructure (EKS, Terraform, networking) that supports AI and enterprise growth. Requires 10+ years infrastructure experience, director-level people leadership of distributed teams, deep architectural expertise, and strong cross-functional influence.
About the job
Responsibilities
- Lead ClickUp's Cloud Platform team to deliver reliable, cost-efficient, and scalable AWS infrastructure that stays 3–6 months ahead of product and enterprise demand.
- Bridge infrastructure and product engineering by building partnerships across product, AI, and enterprise teams to share ownership and maintain velocity.
- Set a forward-looking cloud vision, proactively align with stakeholders, and ensure the platform scales for multi-shard, multi-region growth while meeting security and compliance commitments (M365, BCDR, SOC 2).
- Drive millions in annual savings by migrating data workloads to EKS and self-hosting OpenSearch on EKS.
- Deprecate legacy frontend ALBs and consolidate to a single EKS-managed ALB per shard to enable faster deployments.
- Eliminate single points of ownership across Networking, OpenSearch, and Terraform through cross-training and targeted hiring.
- Stand up recurring alignment cadence with Product, AI, Enterprise, and Security teams.
- Ship automated shard buildout via Backstage to remove manual toil.
- Cut P0/P1 incidents through hardened ingress patterns and Terraform-policy guardrails.
- Deliver Q3/Q4 roadmap covering Agent enablement, centralized IaC, and deployment rollout acceleration tied to AI and Enterprise revenue goals.
- Manage a senior-heavy team distributed across the US, Canada, and Europe (FTEs and contractors) with focus on succession planning and bench depth.
- Own multi-million dollar AWS spend accountability and contractor oversight with a savings-pipeline mindset.
Requirements
- Deep AWS infrastructure mastery: EKS, networking (VPC/TGW/DNS), Terraform at scale with strong architectural opinions.
- Track record setting 6-12 month cloud roadmaps tied to business outcomes (enterprise, AI, cost).
- Cross-team influence across Product, Security, SRE, and DB orgs; ability to break "hesitant to cross the boundary" culture.
- Application-level curiosity and understanding of what the product does.
- People leadership at scale: managed and grown senior-heavy, geographically distributed teams with FTE/contractor blends; build bench depth and succession plans.
- Strong operational rigor: on-call instincts, incident ownership, RCA follow-through.
- Security-first instincts with verify-first mindset (auth enforcement, network isolation, shard boundaries).
- AI & automation mindset for DevSecOps agents, AI-assisted IaC migrations, and skills-based automation.
- BS/MS in CS, Engineering, or equivalent.
- 10+ years in infrastructure/platform engineering with significant time in a director-level role managing distributed teams.
Nice-to-Haves
- AI/ML infrastructure experience: scaling workloads like GPU orchestration, model-serving infra, or cost-optimized compute for inference.
Skills
AWS, EKS, Terraform, Vpc, Tgw, DNS, Opensearch, Backstage, Iac, DevSecOps, Gpu Orchestration, Model Serving
Similar jobs
Engineering Management jobsLeads the technical direction, engineering quality, reliability, and hands-on architecture of a 30-person organization building payment, ledger, wallet, and settlement infrastructure. Requires 10+ years of production software experience, 4+ years leading engineers, and deep distributed-systems and regulated-finance expertise.
Leads a team building foundational security services for Crusoe’s GPU cloud and infrastructure fleet, spanning identity, cryptography, runtime protection, access, vulnerability management, and architectural isolation. Requires 8+ years leading hands-on software or security engineering teams at large-scale infrastructure platforms.
Leads the architecture, production launch, and team buildout for a safety-critical flexible-compute control system that manages AI infrastructure power in coordination with utilities. Requires extensive software engineering and team leadership experience, distributed orchestration expertise, and knowledge of data-center power systems and grid programs.
Leads the product and engineering strategy for AI-native enterprise applications across Finance, HR, and Legal, partnering with executives to transform workflows into intelligent products. Requires 15+ years in enterprise technology, product management, or software engineering and substantial multidisciplinary leadership experience.
Leads the AI Platform organization, setting technical strategy for production AI agents and the shared platform used across the company. Requires 12+ years of engineering experience, including leadership through managers and strong expertise in agent architecture, evaluation, reliability, and guardrails.