Staff Network Production Engineer, Network Ops
Operates and scales Crusoe Cloud’s global edge, backbone, and data center networks for GPU-based HPC infrastructure. The role requires 10+ years of production network operations experience, advanced routing knowledge, automation skills, and participation in 24/7 on-call support.
About the job
Responsibilities
- Implement and operate the global edge, backbone, and data center network for high-performance compute (HPC) clusters with GPUs.
- Monitor network performance and perform advanced troubleshooting and root-cause analysis for incidents; guide post-mortem reviews and improvements.
- Execute network changes across data centers, backbone, and edge infrastructure to implement next-generation designs, scale capacity, and improve network reliability.
- Manage, optimize, and address immediate needs versus long-term goals for the entire network, including frontend, backend, backbone, edge, and public-cloud connectivity.
- Collaborate with Network Engineering and cross-functional teams to ensure the network meets business needs.
- Lead operational-excellence initiatives by developing monitoring, alerting, and systems that ensure high network availability.
- Mentor network engineers and establish best practices for incident response, documentation, and operational readiness.
- Manage vendor relationships and contracts related to network services and equipment.
- Provide visibility into network health through detailed metrics, event statistics, and performance analysis.
- Participate in 24/7 on-call support for the network.
Requirements
- 10+ years of related experience operating at scale in a production environment.
- Strong knowledge of TCP/IP, QoS, BGP, OSPF/IS-IS, EVPN, VXLAN, and MPLS-related technologies such as RSVP-TE and LDP.
- Strong understanding of network monitoring protocols and tools, including SNMP, IPFIX, sFlow/NetFlow, and telemetry.
- Familiarity with data center, backbone, and edge network architecture and implementations.
- Experience with scripting, coding, or programming—Python or similar—to automate operational tasks and build tooling.
- Hands-on experience with major network devices from Mellanox, Cisco, Arista, Juniper, and other mainstream vendors.
- Familiarity with commercial switch/router chipsets such as Broadcom and Barefoot.
- Knowledge of public-cloud connectivity options for AWS, Google Cloud, Azure, Alibaba Cloud, and OCI.
- Understanding of IPv6, IPv4, and IPv4/IPv6 coexistence technologies.
- Experience with network observability, flow analytics, and dashboarding tools such as NSG, Kentik, Grafana, Arbor, ThousandEyes, Catchpoint, and Packet Design.
- Bachelor's degree in Computer Science, Information Science, Engineering, Mathematics, or a related field, or equivalent experience based on three or more years of work experience.
Nice to Have
- Familiarity with backend technologies such as InfiniBand (IB), RoCE, or Spectrum-X.
Compensation and Benefits
- Competitive benefits package including pension contributions, private health and dental insurance, income protection, and life assurance.
- Compensation may be paid as salary or hourly and is determined by education, experience, knowledge, skills, abilities, internal equity, and market data.
Skills
TCP/IP, Qos, BGP, Ospf, Evpn, Vxlan, Mpls, Snmp, Python, Ipv6, AWS, Grafana, InfiniBand, Roce, Network Telemetry
Similar jobs
DevOps / SRE jobsLeads the design, development, and operation of Stripe’s large-scale CI and developer productivity systems. Requires 10+ years of hands-on software development, distributed-systems expertise, technical leadership, and mentoring experience.
Leads the technical direction of multi-cloud Kubernetes capacity management and workload placement across Datadog’s large-scale infrastructure. The role requires strong systems programming experience, ideally in Go, cloud infrastructure expertise, and the ability to influence architecture across teams.
The Staff Production Engineer will operate and improve reliable production infrastructure, with a strong focus on data center networking, capacity planning, troubleshooting, automation, and hardware operations. The role requires 5+ years of production engineering experience, networking expertise, and proficiency with infrastructure and CI/CD tools.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.