Skip to content

Network Engineer, Supercomputing

Own and debug multi-thousand-GPU network fabric (RDMA/RoCE, NVLink/NVSwitch) for large-scale AI training and inference. Requires backend language proficiency, large-scale cluster experience, and cross-stack ownership.

About the job

What You’ll Do

  • Reason about and validate GPU network fabric design across our deployments.
  • Debug RDMA / RoCEv2 across different NIC vendors. Diagnose collective failures of production NCCL, PFC/ECN tuning, and congestion control behavior.
  • Own NVLink / NVSwitch interconnect — including fabric manager and IMEX health, link and lane errors, and how the GPU fabric interacts with collectives.
  • Build host-level network instrumentation and use Linux tooling to build dashboards and alerts, not just the bug report.
  • Navigate cross-cloud fabric quirks across providers and triage across the NIC, driver, kernel, switch, and workload boundaries.
  • Drive escalations with cloud-provider networking teams, owning issues end-to-end until they're resolved.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (we use Python or Rust).
  • Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

Preferred qualifications:

  • Fluency with host-level debugging tools on Linux.
  • Strong communication skills, internally and with cloud providers.
  • Extensive experience with at least one of the following:
    • Familiarity with cloud network primitives across at least two cloud providers.
    • Hands-on experience with NVLink / NVSwitch, fabric manager, and IMEX.
    • Statistical rigor in reliability reasoning — comfort reasoning about failure and error rates, distributions, and base rates, and the judgment to separate signal from noise when characterizing a large fabric.
    • A track record of writing tooling that made the next debugging session meaningfully faster.
    • Familiarity with CUDA/NCCL and performance profiling for distributed training and inference.
    • Understanding of deep learning frameworks and their underlying system architectures.

Compensation and Benefits

  • Compensation: $350,000 - $475,000 USD annual salary range.
  • Benefits: Generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
  • Visa sponsorship available.

Skills

Python, Rust, Kubernetes, Slurm, Rdma, Rocev2, Nccl, Nvlink, Nvswitch, CUDA, Linux

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Site Reliability Engineer (SRE)
$350k+/yrOn-siteDevOps / SRE

Site Reliability Engineer drives end-to-end reliability for AI fine-tuning platform Tinker, including SLOs, monitoring, incident response, and multi-tenant GPU scheduling. Requires distributed systems experience, software proficiency for reliability, and production incident handling.

Anthropic

Anthropic

San Francisco, CA

DevOps / AgentOps Engineer, GTM Systems
$320k+/yrHybridDevOps / SRE

Build and operate an AI-first CI/CD and agent-operations platform for Salesforce and custom GTM applications. The role focuses on governed releases, approval workflows, observability, rollback, sandboxing, and SOX-compliant auditability.

Anthropic

Anthropic

San Francisco, CA
Software Engineer, Infrastructure, Interpretability
$320k+/yrHybridDevOps / SRE

Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.

OpenAI

OpenAI

San Francisco, CA

Systems Integration Engineer, Build Systems | Consumer Devices
$293k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.

OpenAI

OpenAI

San Francisco, CA

Network Engineer
$293k+/yrHybridDevOps / SRE

Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.