Director, Cloud Infrastructure
Leads Sanity’s cloud infrastructure, Platform, and SRE strategy, overseeing Kubernetes, cloud foundations, reliability, deployment, observability, cost, and multi-region architecture. The role requires deep technical expertise, experience managing infrastructure leaders, and a track record operating high-scale production systems.
About the job
Responsibilities
- Set the infrastructure strategy for Sanity’s next stage of scale, with a roadmap across cloud infrastructure, reliability, deployment, observability, security, cost, and developer experience.
- Lead teams responsible for GCP projects and clusters, Kubernetes, networking, service discovery, routing, gateways, CI/CD, infrastructure as code, and observability.
- Define ownership between Platform and SRE, enabling product teams to run services effectively with strong tooling, standards, and incident support.
- Raise reliability standards across production systems, including dashboards, alert severity, paging, service ownership, on-call readiness, and incident response.
- Establish clear deployment paths, production-readiness checks, safe rollouts, and automation.
- Partner with Product, Engineering, Security, Support, Sales, and Customer Success on uptime, latency, scale, compliance, trust, and cost.
- Own cloud cost discipline and make infrastructure tradeoffs visible.
- Shape longer-term architecture for multi-region scale, disaster recovery, data residency, and customer trust requirements.
- Hire, coach, and develop infrastructure leaders and engineers.
Requirements
- Experience leading Infrastructure, Platform, SRE, Cloud, or Developer Platform teams in a scaling SaaS, cloud, infrastructure, API, data, or developer-tools company.
- Experience operating high-volume, multi-region production systems with strict uptime expectations, large cloud bills, customer-facing incidents, and trust requirements.
- Experience building or running production platforms with Kubernetes, GCP or AWS, Terraform or similar infrastructure as code, service discovery, networking, API gateways, CDNs, observability, CI/CD, and incident response.
- Technical depth to debate architecture with senior infrastructure engineers and make pragmatic decisions.
- Experience separating central platform ownership from product-team service ownership.
- Experience improving on-call and reliability through systems, standards, and feedback loops.
- Ability to turn complex infrastructure work into an actionable strategy and drive execution.
- Strong communication skills with executives, product leaders, engineers, and customer-facing teams.
- Experience managing managers or senior technical leads.
Benefits and Compensation
- Real infrastructure scale and a clear mandate to change how it works.
- Senior seat in R&D with close partnership across Product, Engineering, Security, and customer-facing teams.
- Skilled, supportive team and flexible, trust-based work environment.
- Global, culturally diverse colleagues and customers.
- Comprehensive health plans and perks.
- Work-life balance that accommodates individual and family needs.
- Competitive stock options and location-based salary.
Skills
Kubernetes, GCP, Amazon Web Services, Terraform, Networking, Service Discovery, Api Gateways, Content Delivery Networks, Observability, Continuous Integration, Continuous Deployment, Infrastructure As Code, Incident Response, Disaster Recovery
Similar jobs
Engineering Management jobsLeads multiple engineering teams and sets the technical strategy for Spotify’s global content catalog platform. The role requires experience managing through engineering leaders, architecting distributed data-intensive systems, modernizing legacy platforms, and driving cross-functional execution.
Leads and scales a regional or domain-based Customer Engineering organization, owning commercial outcomes, team development, technical go-to-market strategy, and executive customer engagement. Requires 10+ years in technical go-to-market leadership, experience managing managers and revenue targets, and strong cloud, networking, or security fluency.
Leads the engineering organization responsible for Cloudflare’s billing platform and quote-to-cash systems, setting strategy, scaling managers, and ensuring correctness, reliability, auditability, and operational excellence. Requires 15+ years of software engineering experience and substantial experience leading multi-team engineering organizations.
Leads and develops an Enterprise Architecture team that supports strategic customer engagements, revenue growth, and enterprise modernization. The role requires 20+ years of experience, executive communication skills, deep architecture and data expertise, and an MSc or equivalent technology degree.
Leads the Review Tooling team in building investigation, enforcement, analytics, and privacy-aware platforms that help human reviewers and Claude scale safety workflows. Requires engineering management experience, a full-stack or platform background, and cross-functional work with policy, operations, legal, and data science teams.