Staff Network Engineer, Operations
Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
About the job
Responsibilities
- Own production reliability and uptime across global edge, backbone, data center, and GPU cluster networks supporting AI workloads.
- Lead and contribute to high-severity network incident response, including mitigation, stakeholder communication, and postmortem documentation.
- Drive root-cause analyses, identify systemic issues, and track remediation plans through closure.
- Improve network observability using streaming telemetry, SNMP, NetFlow, and monitoring platforms.
- Author and maintain runbooks, escalation playbooks, and standard operating procedures.
- Build Python tooling to automate remediation, diagnostics, and common operational workflows.
- Partner with Architecture and SRE teams to define and track network SLIs and SLOs using real-time dashboards.
- Mentor senior engineers and promote operational excellence and continuous learning.
Requirements
- 8+ years of production network engineering experience focused on operations, incident response, and reliability in large-scale or internet-scale environments.
- Hands-on experience with streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, and ThousandEyes.
- Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads, including PFC, ECN, and DCQCN tuning.
- Expert knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in production data center environments.
- Proficiency with Arista EOS and Juniper Junos platforms in leaf-spine CLOS architectures across multi-vendor environments.
- Python proficiency for auto-remediation scripts, diagnostic tooling, and operational automation.
- Experience operating large device fleets across multiple regions with on-call responsibility and critical-event escalation experience.
- Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience.
Nice-to-haves
- Experience with NVIDIA/Mellanox networking platforms in GPU cluster environments.
- Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility.
- Experience defining or contributing to SLIs and SLOs with SRE or product teams.
- Experience operating fleets of 10,000+ devices in hyperscale or cloud environments.
- Experience contributing to post-incident learning programs or organization-wide operational excellence initiatives.
Compensation and Benefits
- Compensation range of $195,000–$235,000 plus bonus.
- Restricted Stock Units included in all offers.
- Paid time off, paid holidays, and leave programs.
- Comprehensive health, dental, and vision insurance.
- Employer HSA contributions.
- Paid parental leave, life insurance, and short- and long-term disability coverage.
- Professional development and tuition reimbursement.
- Mental health and wellness support.
- Commuter benefits and cell phone stipend.
- 401(k) retirement plan with company match up to 4% of salary.
- Volunteer time off, global travel insurance, emergency assistance, daily meals allowance, and location-specific programs.
Skills
Python, BGP, Evpn-Vxlan, Is-Is, Ospf, Mpls, Qos, TCP/IP, Arista Eos, Juniper Junos, Rdma/Roce, Grafana, Prometheus, Thousandeyes, Snmp
Similar jobs
DevOps / SRE jobsLeads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.
Staff DevSecOps Engineer designing and automating security controls across AWS infrastructure, containers, CI/CD, and platform services. Requires 7+ years of related experience plus expertise in cloud security, infrastructure as code, hardened images, vulnerability scanning, identity, and secrets management.