THE SILICON DISSECTION

What Lives Inside an $30,000 AI Chip?

We know GPUs are built for massive parallel throughput. But when you crack open an 80-billion transistor NVIDIA H100 die, what makes it 10x faster for neural networks than any conventional chip?

โš™๏ธ

CUDA Cores

The general factory workers doing standard single-number math.

14,592 Cores
โšก

Tensor Cores

The hydraulic stamping press that crushes entire matrices in 1 tick.

456 Tensor Cores
๐ŸŒŠ

HBM3e VRAM

The multi-terabyte memory firehose preventing compute starvation.

3,350 GB/s Bandwidth
๐Ÿ”ฌ Let's put on our safety goggles and zoom down to the nanometer level.
THE MODULAR BUILDING BLOCK

The SM: The GPU's Modular City

You don't need to memorize 14,000 cores. You only need to understand one modular neighborhood: the Streaming Multiprocessor (SM).

๐Ÿ™๏ธ
132 Identical Neighborhoods: An NVIDIA H100 chip is simply 132 SMs tiled across the silicon die.
๐Ÿ“ฆ
Self-Contained Powerhouse: Every SM contains its own instruction cache, warp schedulers, register files, and shared L1 memory.
๐Ÿš€
Warp Dispatch: Threads execute in lockstep groups of 32 threads called a Warp. All 32 execute the identical instruction in unison (SIMT).
STREAMING MULTIPROCESSOR (SM) ANATOMY 1 OF 132 SM TILES
Warp Scheduler
16ร— FP32
Tensor Core
Warp Scheduler
16ร— FP32
Tensor Core
Warp Scheduler
16ร— FP32
Tensor Core
Warp Scheduler
16ร— FP32
Tensor Core
256 KB ULTRA-FAST SHARED L1 CACHE & REGISTER FILE
Key Takeaway: Scale AI workloads by distributing blocks of threads across dozens of SMs.
GENERAL PURPOSE ARITHMETIC

CUDA Cores: The Factory Workers

A CUDA core is like a diligent factory worker sitting at a desk with a pocket calculator. Extremely dependable, but calculates one number at a time.

๐Ÿงฎ
Fused Multiply-Add (FMA): Every clock cycle, it computes:
result = A ร— B + C
๐ŸŽฏ
Versatile Generalist: Performs physics simulation, pixel shading, encryption, audio DSP, and ray tracing.
โณ
The Limitation for AI: If you need to multiply a 4ร—4 matrix (16 numbers), 1 CUDA core requires 16 clock cycles.
CUDA CORE FMA PIPELINE 1 OP / CLOCK CYCLE
Number A
ร—
Number B
+
Number C
FP32 / INT32 ARITHMETIC LOGIC UNIT (ALU)
Output = 1 Single Number (D)
The Bottleneck: 70 billion parameters take too long when computed one scalar number at a time.
NVIDIA'S SECRET WEAPON

Tensor Cores: The 4ร—4 Matrix Stamping Press

Instead of calculating scalar numbers, a Tensor Core consumes two whole 4ร—4 matrices and outputs the full matrix product in 1 single clock cycle!

CUDA Core (Scalar FMA) Cycles: 0 / 16
Tensor Core: 16x faster!
Tensor Core (Matrix MMA) Cycles: 0 / 1
๐Ÿ’ฅ

Hardware Acceleration for GEMM: $D = A \times B + C$. Tensor Cores physically hardwire the matrix dot product into silicon, generating the 10x throughput leap that made Modern LLMs possible!

THE MEMORY WALL

VRAM & HBM: The Bandwidth Firehose

You can have the fastest Tensor Cores on planet Earth, but if they run out of numbers to crunch, your $30,000 GPU sits completely idle.

๐Ÿฅค
DDR5 = The Coffee Straw (300 GB/s): Standard CPU memory travels through motherboard bus traces. Too slow to feed 14,000 GPU cores.
๐ŸŒŠ
HBM3e = Niagara Falls (3,350 GB/s): 3D-stacked silicon dies sit on the exact same substrate as the GPU, providing a 10x wider data pipeline.
โš–๏ธ
Bandwidth vs Capacity: An H100 has 80GB VRAM, but what matters most is how fast that 80GB can be streamed into the compute registers!
MEMORY BANDWIDTH COMPARISON HBM3e vs DDR5
Server CPU (DDR5) 300 GB/s
Gaming GPU (GDDR6X - RTX 4090) 1,008 GB/s
Enterprise AI GPU (HBM3e - H100) 3,350 GB/s (11x DDR5)
The Memory Wall: When an LLM runs inference, speed is bounded by memory bandwidth, not raw compute!
PRECISION REVOLUTION

Precision Wars: FP32 vs. FP16 vs. FP8

Why neural networks don't need 32-bit perfectionโ€”and how lowering precision 4x'd AI speeds.

FP32 (Single Precision) 4 Bytes / number
1b 8b Exponent 23b Mantissa (Fraction)

Pinpoint scientific accuracy. Traditional HPC default. Consumes 4x memory.

Base Speed (1x)
FP16 / BF16 (Half) 2 Bytes / number
1b 5b / 8b 10b / 7b Mantissa

The modern AI training standard. Halves VRAM usage with zero loss in model quality.

2x Faster Compute
FP8 (Quarter Precision) 1 Byte / number
1b 4b Exponent 3b Mantissa

Hopper & Blackwell breakthrough. Fits 70B models in 1/4 the memory space!

4x Faster Compute (TFLOPS)
CHAPTER 02 SUMMARY

Inside the GPU: Quick Review

1

SMs are the Modular Cities

132 identical multiprocessor tiles scale parallel workload distribution.

2

CUDA vs Tensor Cores

CUDA cores do 1 number at a time; Tensor cores stamp entire 4ร—4 matrices in 1 tick.

3

HBM Memory Solves the Memory Wall

3.35 Terabytes/sec data firehose keeps 14,000 parallel cores saturated.

UP NEXT: CHAPTER 03

GPU Memory: Why VRAM Matters for LLMs

Ever seen the dreaded error: CUDA Out of Memory: tried to allocate 2.4 GiB? Where does your 80GB VRAM actually go between Model Weights, KV Cache, and Activations?

โš ๏ธ torch.cuda.OutOfMemoryError: CUDA out of memory.