HIGH-BANDWIDTH INTRA-NODE FABRIC

NVIDIA NVLink: How 8 GPUs Become One

Inside the 900 GB/s interconnect that turns eight individual chips into a 640GB unified silicon brain.

THE STANDARD PCIE GEN5 TRAP

Standard PCIe Bus (Two-Lane Country Road)

64 GB/s
Unidirectional Peak Throughput
GPU 0 → ⚠️ CONGESTION → Host CPU / PCIe Switch → ⚠️ 125ms DELAY → GPU 7

When 8 GPUs synchronize attention heads after every layer, PCIe saturates instantly, cutting Tensor Core efficiency to ~30%.

The NVLink Generations: Scaling the Wire

From the external dual-GPU SLI bridges of the 2000s to dense PCB traces capable of moving terabytes per second.

NVLink 1 (Pascal) 2016
160 GB/s

First high-speed bridge on P100. 5x faster than PCIe Gen3.

Direct PCB traces
NVLink 2 (Volta) 2018
300 GB/s

6 links on V100. Introduced the first physical NVSwitch chip.

First NVSwitch generation
NVLink 3 (Ampere) 2020
600 GB/s

12 links on A100 SXM4. Made Tensor Parallelism mainstream.

SXM4 Integrated Backplane
NVLink 4 (Hopper) 2022
900 GB/s

18 links on H100 SXM5. Hardware asynchronous FP8 memory copies.

14x faster than PCIe Gen5
NVLink 5 (Blackwell) 2024
1,800 GB/s

1.8 TB/s bidirectional on B200. Powers the 72-GPU NVL72 rack.

Rack-Scale 130 TB/s Fabric

Inside HGX: The NVSwitch Crossbar Matrix

Without NVSwitch, sending data across 8 GPUs requires multi-hop daisy-chaining. NVSwitch creates an instant zero-hop crossbar.

GPU 0
GPU 1
GPU 2
GPU 3
NVSwitch 1
NVSwitch 2
NVSwitch 3
NVSwitch 4
FULL NON-BLOCKING CROSSBAR (7.2 TB/s AGGREGATE)
GPU 4
GPU 5
GPU 6
GPU 7
🎯
Zero Hop Penalty: Any GPU communicates with any other GPU in exactly 1 direct switch hop.
Hardware Multicast: One GPU can broadcast a tensor to all 7 peers simultaneously without CPU copying.
🧠
SHARP Reduction inside NVSwitch: Gradient math performed inside the internal crossbar chips!

Interconnect Speed & Attention Sync Simulator

Simulate an 8-way Tensor Parallel All-Gather step. Compare standard PCIe Gen5 vs NVLink 4 vs Blackwell NVLink 5!

8 GB
All-Gather attention activations across all 8 GPUs per transformer layer.
NVIDIA HGX H100 (8x SXM5 + 4x NVSWITCH) READY // NVLINK 4
Sync Transfer Duration
8.8 ms
Blazing fast All-Gather
Effective Compute MFU
89%
Tensor Cores rarely wait
Interconnect Bandwidth
900 GB/s
Per-GPU Bidirectional
Remote Memory Latency
280 ns
Hardware load/store speed

Memory Pooling: The 640GB Virtual Mega-GPU

NVLink doesn't just transfer packets. It merges 8 separate HBM3 memory banks into a single coherent virtual memory space.

8 PHYSICAL HBM3 POOLS (80GB EACH)
GPU 0: 80GB
GPU 1: 80GB
GPU 2: 80GB
GPU 3: 80GB
GPU 4: 80GB
GPU 5: 80GB
GPU 6: 80GB
GPU 7: 80GB
NVSWITCH CACHE-COHERENT FABRIC // SUB-300ns DIRECT ACCESS
LOGICAL VIEW (PYTORCH / VLLM)

One Coherent 640GB Virtual Mega-GPU

// GPU 0 executes standard assembly load straight from GPU 7's memory! LD.E.64 R2, [remote_gpu7_hbm_address]; // Completes in 280 nanoseconds!
✓ Zero OS Kernel Socket Copies
✓ Zero Network Protocol Overhead
✓ 70B & 405B Models Fit in Unified Memory

Rack-Scale NVLink: The GB200 NVL72

Blackwell breaks out of the 8-GPU chassis. 5,000 direct-drive copper cables turn 72 GPUs into one unified mega-system.

GB200 NVL72 RACK SPECIFICATIONS
Total Compute: 72 x Blackwell GPUs (36 Dual-Grace Nodes)
Aggregate NVLink Bandwidth: 130 Terabytes / Second
Unified Fast Memory: 30 Terabytes (13.5 TB HBM3e + 16.2 TB LPDDR5X)
Interconnect Spine: 5,000 Direct-Drive Copper Cables (2 Miles of Wire!)
Power Consumption: 120 kW (100% Liquid Cooled)
5,000 DIRECT-DRIVE PASSIVE COPPER CABLES
Why Passive Copper? At 1.8 TB/s per GPU, optical transceivers would consume 20 kilowatts of power just for lasers! NVIDIA engineered direct-drive passive copper cartridges to save 20kW of power and eliminate optical failures.

Three Core Truths of NVIDIA NVLink

Remember these fundamentals when evaluating AI servers and multi-GPU architectures.

01

PCIe is a Barrier

Standard PCIe Gen5 (64 GB/s) bottlenecks Tensor Parallelism. NVLink 4 (900 GB/s) is 14x faster, keeping Tensor Cores fed at 89%+ MFU.

900 GB/s Per GPU
02

NVSwitch Crossbar

4 NVSwitch chips eliminate daisy-chain hops, giving any GPU full bisection bandwidth to any other GPU with zero intermediate delays.

7.2 TB/s Chassis Fabric
03

Unified Memory Pooling

8 physical GPUs become one 640GB virtual mega-chip with hardware load/store instructions completing in under 300 nanoseconds.

Cache-Coherent NUMA
NEXT IN CHAPTER 09

Scaling from 1 GPU to 1,000 GPUs: 3D Parallelism

How to train massive models when they outgrow a single server: Data Parallelism (DP), Tensor Parallelism (TP), Pipeline Parallelism (PP), and ZeRO Stage 3!

Coming Up Next