Senior Network Production Operations Engineer
Operates and scales Crusoe Cloud’s global edge, backbone, and data center networks supporting GPU-based HPC workloads. The role requires extensive production networking experience, strong protocol and observability expertise, automation skills, and participation in 24/7 on-call support.
About the job
Responsibilities
- Monitor network performance and perform advanced troubleshooting and root-cause analysis for incidents.
- Guide post-mortem reviews and operational improvements.
- Execute network changes across data centers, backbone, and edge infrastructure.
- Manage and optimize the global network across frontend, backend, backbone, edge, and public-cloud connectivity.
- Collaborate with Network Engineering and cross-functional teams.
- Develop monitoring, alerting, and systems to ensure high network availability.
- Mentor network engineers and establish best practices for incident response, documentation, and operational readiness.
- Manage vendor relationships and contracts related to network services and equipment.
- Deliver network-health metrics, event statistics, and performance analysis.
- Participate in 24/7 on-call support.
Requirements
- 10 years of related experience operating production environments at scale.
- Strong knowledge of TCP/IP, BGP, OSPF/IS-IS, EVPN, VXLAN, QoS, MPLS, RSVP-TE, and LDP.
- Knowledge of network monitoring protocols and tools, including SNMP, IPFIX, sFlow/NetFlow, and telemetry.
- Familiarity with data center, backbone, and edge network architecture.
- Scripting or programming experience, preferably Python, for automation and tooling.
- Hands-on experience with Mellanox, Cisco, Arista, Juniper, and other mainstream network vendors.
- Familiarity with commercial switch/router chipsets such as Broadcom and Barefoot.
- Knowledge of connectivity options for AWS, Google Cloud, Azure, Alibaba Cloud, and Oracle Cloud Infrastructure.
- Understanding of IPv4, IPv6, and coexistence technologies.
- Experience with network observability, flow analytics, and dashboarding tools such as NSG, Kentik, Grafana, Arbor, ThousandEyes, Catchpoint, and Packet Design.
- Bachelor's degree in Computer Science, Information Science, Engineering, Mathematics, or a related field, or equivalent experience.
Nice to Have
- Familiarity with InfiniBand, RoCE, or Spectrum-X.
Compensation & Benefits
- Competitive benefits package including pension contributions, private health and dental insurance, income protection, and life assurance.
- Compensation may be paid as salary or hourly and is determined based on education, experience, skills, internal equity, and market data.
Skills
TCP/IP, BGP, Ospf, Is-Is, Evpn, Vxlan, Mpls, Snmp, Ipfix, Netflow, Python, Cisco, Arista, Juniper, AWS
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.
Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.