Engineering Manager, Infrastructure
Leads infrastructure team managing cloud, networking, storage, and compute for high-scale developer tool. Sets technical direction, codes, hires, and optimizes costs/regional deployments with deep Kubernetes/AWS expertise.
About the job
Responsibilities
- Lead the infrastructure team owning cloud, networking, storage, and compute layers including network foundations, container orchestration, edge and security infrastructure, data storage systems, and compute runtimes.
- Set technical direction, write and review code, and manage a team of infrastructure engineers.
- Drive cost management, regional deployment strategy, and infrastructure unification.
- Example projects: Own Kubernetes clusters with service mesh, scaling, and ingress; design geo-deployment architecture; build edge and security infrastructure; own data storage strategy (Postgres, OLAP, caching); manage cost optimization; unify compute platform; hire and grow the team.
Requirements
- Led engineering teams building/operating production infrastructure or platform systems at scale.
- Deep experience with AWS (VPC networking, EKS/K8s, IAM/account management) or comparable clouds.
- Built/operated production Kubernetes clusters at scale (service mesh, autoscaling, multi-region).
- Strong knowledge of databases, storage engines, caching, schema design, and tradeoffs (performance, consistency, cost).
- Experience with edge networking, CDN/WAF, traffic management.
- Care about infrastructure-as-code, reproducibility, self-service for teams.
Nice-to-Haves
- Cost optimization at scale, infrastructure migration/unification, data storage systems (Postgres, ClickHouse, OLAP).
Skills
Kubernetes, AWS, EKS, Postgres, Service Mesh, Vpc, IAM, ClickHouse, Olap, Infrastructure As Code
Similar jobs
DevOps / SRE jobsDesigns and operates foundational developer-infrastructure services for CI, builds, deployments, and testing. The role requires senior-level systems engineering, end-to-end service ownership, and cross-functional technical leadership.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.