Staff Network Engineer, Deployment
Leads physical and logical deployment of global network infrastructure for AI data centers, including rack/stack, cabling, automation with Python/Ansible, testing, and partner coordination. Requires 8+ years experience with Arista, Juniper, Mellanox, BGP/EVPN, and physical layer expertise.
About the job
What You’ll Be Working On
- Execute Global Build-outs: Lead the end-to-end deployment of network infrastructure in new and existing data centers, from initial rack/stack oversight to final hand-off.
- Bridge Design and Reality: Take high-level designs from the Network Development team and translate them into site-specific implementation plans, cable maps, and configuration templates.
- Validate and Commission: Perform rigorous "Burn-in" testing and site acceptance testing (SAT) for new network clusters, ensuring zero-defect handovers to the Operations team.
- Optimize Deployment Automation: Use Python, Ansible, and ZTP (Zero Touch Provisioning) to automate the staging and configuration of hundreds of network devices simultaneously.
- Manage On-site Partners: Coordinate with remote hands, structured cabling vendors, and data center providers to ensure physical layer standards (fiber paths, power requirements, and cooling) meet Crusoe’s stringent HPC requirements.
- Inventory and Capacity Management: Track global hardware assets and lead the "Turn-up" of new backbone capacity and edge interconnects.
What You’ll Bring to the Team
- 8+ years of experience in network engineering with a heavy focus on large-scale data center deployments and infrastructure projects.
- Mastery of Physical Layer Standards: Expert knowledge of structured cabling (SMF/MMF, MPO/MTP), optical transceivers (400G/800G), and data center power/cooling requirements.
- Strong Routing and Switching Knowledge: Hands-on experience configuring Arista (EOS), Juniper (Junos), and NVIDIA/Mellanox platforms in a leaf-spine architecture.
- Protocol Proficiency: Solid understanding of BGP, EVPN-VXLAN, and LLDP as they relate to large-scale fabric provisioning.
- Automation-First Mindset: Proficiency in Python and Ansible for automating repetitive deployment tasks and validating configuration state.
- Logistical Excellence: Proven ability to manage multiple complex projects simultaneously across different time zones and physical locations.
- Troubleshooting Expertise: Ability to diagnose complex physical layer and link-layer issues using OTDRs, light meters, and packet captures.
- Education: Bachelor’s degree in a technical field or equivalent practical experience in hyperscale or ISP environments.
Compensation
$193,000 - $234,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant’s education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data.
Skills
Arista Eos, Juniper Junos, Nvidia Mellanox, BGP, Evpn-Vxlan, Lldp, Python, Ansible, Zero Touch Provisioning, Structured Cabling, Smf, Mmf, Mpo, Optical Transceivers
Similar jobs
DevOps / SRE jobsOwn reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.