NVIDIA NVLink: How 8 GPUs Become One
Inside the 900 GB/s interconnect that turns eight individual chips into a 640GB unified silicon brain.
Standard PCIe Bus (Two-Lane Country Road)
When 8 GPUs synchronize attention heads after every layer, PCIe saturates instantly, cutting Tensor Core efficiency to ~30%.
The 900 GB/s Optical/Copper Highway
Zero CPU hops. 18 high-speed links connect every GPU directly into a non-blocking NVSwitch crossbar in under 300 nanoseconds.
The NVLink Generations: Scaling the Wire
From the external dual-GPU SLI bridges of the 2000s to dense PCB traces capable of moving terabytes per second.
First high-speed bridge on P100. 5x faster than PCIe Gen3.
6 links on V100. Introduced the first physical NVSwitch chip.
12 links on A100 SXM4. Made Tensor Parallelism mainstream.
18 links on H100 SXM5. Hardware asynchronous FP8 memory copies.
1.8 TB/s bidirectional on B200. Powers the 72-GPU NVL72 rack.
Inside HGX: The NVSwitch Crossbar Matrix
Without NVSwitch, sending data across 8 GPUs requires multi-hop daisy-chaining. NVSwitch creates an instant zero-hop crossbar.
Interconnect Speed & Attention Sync Simulator
Simulate an 8-way Tensor Parallel All-Gather step. Compare standard PCIe Gen5 vs NVLink 4 vs Blackwell NVLink 5!
Memory Pooling: The 640GB Virtual Mega-GPU
NVLink doesn't just transfer packets. It merges 8 separate HBM3 memory banks into a single coherent virtual memory space.
One Coherent 640GB Virtual Mega-GPU
// GPU 0 executes standard assembly load straight from GPU 7's memory!
LD.E.64 R2, [remote_gpu7_hbm_address]; // Completes in 280 nanoseconds!
Rack-Scale NVLink: The GB200 NVL72
Blackwell breaks out of the 8-GPU chassis. 5,000 direct-drive copper cables turn 72 GPUs into one unified mega-system.
Three Core Truths of NVIDIA NVLink
Remember these fundamentals when evaluating AI servers and multi-GPU architectures.
PCIe is a Barrier
Standard PCIe Gen5 (64 GB/s) bottlenecks Tensor Parallelism. NVLink 4 (900 GB/s) is 14x faster, keeping Tensor Cores fed at 89%+ MFU.
NVSwitch Crossbar
4 NVSwitch chips eliminate daisy-chain hops, giving any GPU full bisection bandwidth to any other GPU with zero intermediate delays.
Unified Memory Pooling
8 physical GPUs become one 640GB virtual mega-chip with hardware load/store instructions completing in under 300 nanoseconds.