Data Center Provisioning Engineer
Provisions, commissions, and validates network, server, and related infrastructure across large-scale data center deployments. The role requires 5+ years of relevant experience, strong networking and systems troubleshooting skills, and proficiency in automation, cloud platforms, Kubernetes, Terraform, and Ansible.
About the job
Responsibilities
- Provision and configure network devices, firewalls, and servers in data centers using documented processes, provisioning frameworks, and automation tools.
- Troubleshoot server, network, configuration, automation, and connectivity issues during infrastructure provisioning and service bring-up.
- Coordinate with engineering teams, data center technicians, cabling vendors, and rack integrators to support infrastructure deployment and site readiness.
- Support network, power, and mechanical commissioning, including testing and validation of connectivity, device configurations, server readiness, and infrastructure dependencies.
- Improve provisioning processes, tooling, and automation for reliable, repeatable, and scalable deployment across multiple sites.
- Perform post-deployment validation and integrate infrastructure with monitoring, alerting, and operational tooling.
- Coordinate handoff of deployed infrastructure to Cluster Operations, Network Operations, and Site Operations teams.
- Document deployment procedures, troubleshooting findings, lessons learned, and process improvements.
Requirements
- Bachelor's degree in Electrical Engineering, Computer Science, Computer Engineering, or a related technical field.
- At least 5 years of relevant experience in network device and server provisioning, data center infrastructure, Site Reliability Engineering, or a related field.
- Hands-on experience provisioning and troubleshooting servers and network devices across compute, networking, operating systems, configuration, and automation layers.
- Proficiency in Python and shell scripting, including automation for infrastructure deployment.
- Experience with at least one major cloud platform: AWS, Google Cloud, or Azure.
- Experience with Kubernetes, including deployment, scaling, troubleshooting, and cluster management.
- Experience with Infrastructure as Code and configuration management, particularly Terraform and Ansible.
- Strong knowledge of PXE, DNS, DHCP, TCP/IP, load balancing, VPCs, firewalls, VLANs, LACP, BGP, and OSPF.
- Experience upgrading server firmware, switch operating system software, and PDU firmware.
- Strong problem-solving, communication, collaboration, and documentation skills.
Compensation and Benefits
- Work on a breakthrough AI platform and high-performance AI supercomputers.
- Opportunities to publish and open-source AI research.
- Startup vitality with job stability.
- Non-corporate work culture with respect for individual beliefs.
- Inclusive environment with continuous learning, growth, and team support.
Skills
Python, Shell Scripting, AWS, GCP, Azure, Kubernetes, Terraform, Ansible, Pxe, DNS, Dhcp, TCP/IP, BGP, Ospf, Vlans
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.