Skip to content
xAIxAI

Network Development Engineer, ML Infrastructure (High-Speed Interconnects)

Designs, builds, and optimizes high-speed copper and optical interconnects for large-scale AI/ML clusters. Requires 8+ years experience in high-speed networking, deep knowledge of SerDes, photonics, and Master's/PhD in EE/Photonics/Physics.

About the job

Responsibilities

  • Design, validate, and productize high-speed copper and optical connectivity solutions for AI clusters (100k+ GPU scale).
  • Own vendor due diligence and onboarding for new 1.6T products including AEC and pluggable optical transceivers (DR4/8, FR4) including rigorous bring-up & characterization.
  • Investigate the opportunity for LPO and LRO in our network.
  • Evaluate early co-packaged and near-packaged engines for switches and GPUs.
  • Pathfinding for new interconnect modalities including VCSEL, microLED, THz radio-based solutions to improve network economics and reliability.
  • Work closely with vendors (transceiver, cable, SerDes, DSP, silicon photonics foundries) to influence roadmaps and ensure timely delivery of next-gen solutions.
  • Collaborate with ML training teams to translate workload communication patterns into concrete interconnect topology and optical reconfigurability requirements.
  • Perform system-level simulation of end-to-end fabric performance.
  • Drive failure analysis, root cause, and corrective actions for interconnect-related issues in production clusters through fleet-level metrics gathering and analysis.
  • Contribute to internal tooling and automation for interconnect health monitoring, telemetry, diagnostics, remediation and automated qualification pipelines.
  • Stay current with industry standards (OIF CMIS, IEEE) and emerging technologies (multi-core/hollow-core fiber, 448G SerDes, TFLN, ring resonators).

Qualifications

  • At least 8+ years of hands-on experience in designing, deploying and operating high-speed copper and optical interconnects, preferably in a module design role or in a hyperscale datacenter environment.
  • Master's or PhD degree in Electrical Engineering, Photonics or Physics.
  • Deep knowledge of PAM4 SerDes performance, equalization, jitter, crosstalk.
  • Solid operational understanding of FEC, Retimers, TIAs and Drivers.
  • Deep knowledge of optical link budget analysis and performance metrics including TDECQ, OMA, Tcode, stressed receiver sensitivity and associated diagnostics.
  • Expertise in transceiver components including CW lasers, SiPh PICs, EML, DSP, passive subassemblies, their failure modes and characterization.
  • Knowledge of thermal, mechanical, power, signal integrity constraints in dense hardware.
  • Knowledge of SiPh design process, yield improvement and reliability testing.
  • Familiarity with CPO technologies and challenges/risk areas.
  • Familiarity with subcomponent supply chains and global manufacturers, ODMs and CMs.
  • Strong problem-solving skills and ability to thrive in a fast-paced, ambiguous setting.

Compensation and Benefits

Annual Base Salary: $180,000 - $440,000 USD

Base salary is just one part of total rewards package, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

Skills

Pam4 Serdes, Optical Interconnects, Copper Interconnects, Silicon Photonics, Fec, Retimers, Tia, Driver, Tdecq, Oma, Siph Pics, Eml, Dsp, Cpo, Oif Cmis

Runpod

Runpod

United States

Senior HPC Storage Engineer
$180k+/yrRemote8+ YOEDevOps / SRE

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.

Onebrief

Onebrief

Colorado Springs, CO

Senior Site Reliability Engineer, Colorado Springs
$180k+/yrOn-site5+ YOEDevOps / SRE

Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.