Software Engineer, Infrastructure
Builds and operates core infrastructure including Kubernetes clusters, geo-deployments, edge security, and cost optimization to support rapid growth and team productivity. Requires deep AWS/K8s experience and strong engineering fundamentals.
About the job
Responsibilities
- Own Kubernetes and cluster foundations: build and operate production clusters with service mesh, scaling, and ingress.
- Design geo-deployment architecture: build replicable process for deploying geo-replicated services across cloud regions and providers.
- Build edge and security infrastructure: design networking and security layer at edge for abuse protection, rate limiting, and traffic routing.
- Own cost management and optimization: build attribution systems, identify waste, and ensure smart tradeoffs between cost and reliability.
- Unify compute platform: define single container orchestration strategy for consistent, reliable deployments.
Requirements
- Deep experience with AWS (or comparable), especially VPC networking, EKS/K8s, and IAM/account management.
- Built and operated production Kubernetes clusters at scale, including service mesh, autoscaling, and multi-region deployments.
- Understand edge networking, CDN/WAF architectures, and traffic management at infrastructure level.
- Care about infrastructure-as-code, reproducibility, and enabling self-serve reliable infrastructure for other teams.
- Strong software engineering fundamentals.
Nice-to-Haves
- Experience with cost optimization at scale or infrastructure migration/unification.
Skills
Kubernetes, AWS, EKS, Vpc, IAM, Service Mesh, Cdn, Waf, Infrastructure As Code, Autoscaling
Similar jobs
DevOps / SRE jobsDesigns and operates foundational developer-infrastructure services for CI, builds, deployments, and testing. The role requires senior-level systems engineering, end-to-end service ownership, and cross-functional technical leadership.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.