A Curated Guide to Networking Fabrics for AI GPU Clusters
Summary
Unawesome AI Fabric Engineering is a curated reference for GPU performance engineers moving into the networking side of AI clusters. It assumes familiarity with GPU kernels, profiling, inference engines, and distributed inference, and organizes resources from a single NIC-to-GPU path through complete cluster fabrics. The guide covers RDMA verbs, memory registration, NIC behavior, measurement, GPUDirect, PCIe peer-to-peer DMA, GPU-initiated networking, and libraries such as NCCL, RCCL, UCX, and NVSHMEM. It also collects work on collective algorithms, topology-aware communication, in-network reduction, and GPU-driven primitives. For inference systems, it points to resources on transferring state across GPUs, CPUs, storage, and networks, including KV-cache movement, expert routing, and prefill/decode disaggregation. Further sections address congestion control and transports such as DCQCN, TIMELY, HPCC, SRD, Falcon, and Ultra Ethernet, followed by fabric designs for large training clusters, scale-up interconnects, and cloud GPU networks. The operating-at-scale section includes production reports, anomaly detection, validation, fault tolerance, and straggler analysis. The repository separates a frontier section verified on October 5, 2026, and a watchlist for technologies awaiting measured deployments, including Ultra Ethernet systems, adaptive routing, co-packaged optics, UALink, heterogeneous GPU KV-cache transfer, and non-NVIDIA GPU-initiated networking. Its hardware policy excludes software emulation and requires performance claims to identify hardware, topology, message sizes, software versions, and baselines. The collection is therefore intended as a practical, measurement-oriented starting point and reference for real AI networking systems.