Software Engineer — GPU Networking & Distributed Systems
Software engineer architecting GPU networking software with RDMA for distributed AI inference. Optimizes communication for H100/Blackwell clusters, WideEP, and serverless LLM startups using C++/Python and NVIDIA tools.
About the job
What You'll Do
- Make RDMA First-Class: Integrate RDMA/RoCE/InfiniBand into inference stack for bandwidth/latency improvements.
- Optimize Distributed Inference: Implement networking for Disaggregated KV Cache Offload and WideEP across NVLink/InfiniBand.
- Enable Serverless-Grade Startup Speeds: Work on checkpointing/storage for sub-10s startup of trillion-parameter models.
- Deep-Dive into Hardware: Characterize H100/H200, B200/B300, GB200/300 NVL72 clusters with acceptance tests.
- Build Observability: Design tools for packet flow, congestion, bandwidth visualization.
- Optimize Kernels: Use NCCL, NVSHMEM; write custom communication kernels.
Who You Are
- Deep experience with high-performance networking (InfiniBand, RoCE v2).
- Fluent in C++ or Python, understanding NVIDIA memory hierarchy (H100/Blackwell).
- Comfortable diving into TensorRT-LLM source, custom bindings, NVLink debugging.
- Know when to use/build custom solutions beyond standard tools.
Highly Preferred
- Deep knowledge of NCCL, NVSHMEM, UCX.
- Experience with GPUDirect Storage (GDS), Weka/3FS.
- Familiarity with TensorRT-LLM, vLLM, Sglang.
- Running low-level benchmarks for hardware qualification.
Benefits
- Competitive compensation with equity.
- 100% medical/dental/vision coverage.
- Generous PTO, Winter Break.
- Paid parental leave, 401(k).
Skills
Rdma, Roce, InfiniBand, Nccl, Nvshmem, Ucx, C++, Python, Tensorrt-Llm, Nvlink
Similar jobs
Backend Engineering jobsBuild secure, scalable identity and access management platforms, APIs, and internal tooling using JavaScript, Node.js, React.js, and distributed-systems expertise. The role requires 5+ years of backend or distributed-systems experience and familiarity with authentication protocols is a plus.
Build backend systems and APIs that ingest, process, and route large-scale telemetry data in real time. The role requires strong computer science fundamentals, Node.js/TypeScript experience, distributed-systems knowledge, and ownership of reliable cloud software.
Build and maintain highly available, scalable distributed systems, data pipelines, messaging systems, and services powering Internet intelligence products. The role requires 5+ years of distributed-systems experience and proficiency with Go, cloud platforms, databases, and messaging technologies.
Build and maintain backend systems powering Cash App Taxes, translating complex federal and state tax requirements into reliable filing and e-file functionality. The role requires 4+ years of software development experience, strong system design skills, and a bachelor's degree or equivalent experience.
Build shared Ruby backend foundations for APIs, asynchronous jobs, workflow orchestration, and developer tooling. The role requires 5+ years of software engineering experience, including 3+ years with production Ruby systems, plus expertise in distributed systems and large codebases.