Skip to content
CrusoeCrusoe

Staff Network Engineer, Operations

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.

About the job

Responsibilities

  • Own production reliability and uptime across global edge, backbone, data center, and GPU cluster networks supporting AI workloads.
  • Lead and contribute to high-severity network incident response, including mitigation, stakeholder communication, and postmortem documentation.
  • Drive root-cause analyses, identify systemic issues, and track remediation plans through closure.
  • Improve network observability using streaming telemetry, SNMP, NetFlow, and monitoring platforms.
  • Author and maintain runbooks, escalation playbooks, and standard operating procedures.
  • Build Python tooling to automate remediation, diagnostics, and common operational workflows.
  • Partner with Architecture and SRE teams to define and track network SLIs and SLOs using real-time dashboards.
  • Mentor senior engineers and promote operational excellence and continuous learning.

Requirements

  • 8+ years of production network engineering experience focused on operations, incident response, and reliability in large-scale or internet-scale environments.
  • Hands-on experience with streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, and ThousandEyes.
  • Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads, including PFC, ECN, and DCQCN tuning.
  • Expert knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in production data center environments.
  • Proficiency with Arista EOS and Juniper Junos platforms in leaf-spine CLOS architectures across multi-vendor environments.
  • Python proficiency for auto-remediation scripts, diagnostic tooling, and operational automation.
  • Experience operating large device fleets across multiple regions with on-call responsibility and critical-event escalation experience.
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience.

Nice-to-haves

  • Experience with NVIDIA/Mellanox networking platforms in GPU cluster environments.
  • Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility.
  • Experience defining or contributing to SLIs and SLOs with SRE or product teams.
  • Experience operating fleets of 10,000+ devices in hyperscale or cloud environments.
  • Experience contributing to post-incident learning programs or organization-wide operational excellence initiatives.

Compensation and Benefits

  • Compensation range of $195,000–$235,000 plus bonus.
  • Restricted Stock Units included in all offers.
  • Paid time off, paid holidays, and leave programs.
  • Comprehensive health, dental, and vision insurance.
  • Employer HSA contributions.
  • Paid parental leave, life insurance, and short- and long-term disability coverage.
  • Professional development and tuition reimbursement.
  • Mental health and wellness support.
  • Commuter benefits and cell phone stipend.
  • 401(k) retirement plan with company match up to 4% of salary.
  • Volunteer time off, global travel insurance, emergency assistance, daily meals allowance, and location-specific programs.

Skills

Python, BGP, Evpn-Vxlan, Is-Is, Ospf, Mpls, Qos, TCP/IP, Arista Eos, Juniper Junos, Rdma/Roce, Grafana, Prometheus, Thousandeyes, Snmp

Shield AI

Shield AI

San Diego, CA

Senior Staff Lead Site Reliability Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.

Komodo Health

Komodo Health

United States

Staff Infrastructure Engineer
$187k+/yrRemote8+ YOEDevOps / SRE

Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.

VGS

VGS

United States
Senior Staff Infrastructure Engineer
$185k+/yrRemote10+ YOEDevOps / SRE

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.

Shield AI

Shield AI

San Mateo, CA
Staff DevSecOps Engineer
$182k+/yrOn-site7+ YOEDevOps / SRE

Staff DevSecOps Engineer designing and automating security controls across AWS infrastructure, containers, CI/CD, and platform services. Requires 7+ years of related experience plus expertise in cloud security, infrastructure as code, hardened images, vulnerability scanning, identity, and secrets management.