THE $100,000,000 INFRASTRUCTURE QUESTION

Why Did AI Fall in Love With GPUs?

When OpenAI built the clusters to train GPT-4, they didn't buy thousands of 64-core Intel or AMD server CPUs. They spent hundreds of millions of dollars buying tens of thousands of GPUs.

CPU
Core 1
Core 2
Core 3
Core 4
MASSIVE L3 CACHE

64-Core Server CPU

The undisputed king of data centers for 30 years. Runs Linux, databases, and microservices.

❌ Not Chosen for AI Training
VS
GPU
HBM3e HIGH-BANDWIDTH MEMORY

NVIDIA H100 Tensor Core GPU

Originally engineered for rendering 3D video games like Cyberpunk 2077.

πŸ‘‘ The $100B Engine of Generative AI
πŸ€” Why did a chip built for video games become the holy grail of machine learning? Let's break it down in 4 minutes.
THE GENERALIST COMPUTATION ENGINE

Meet the CPU: The Michelin-Star Chef

Imagine the world's most brilliant Master Chef. They can invent complex recipes, taste and adjust delicate seasoning, manage kitchen inventory, and talk to restaurant guests simultaneously.

⚑
Ultra-Low Latency: Blazing clock speeds (4.5 – 5.0+ GHz). Finishes any individual task in nanoseconds.
🧠
Branch Prediction Genius: Handles complicated if/else statements, context switches, and interrupts without breaking a sweat.
⚠️
The Bottleneck: Only 2 Hands! A CPU only has 8 to 64 cores. If you ask the master chef to chop 50,000 onions, they must do it one by one!
CENTRAL PROCESSING UNIT DIE (64-CORE) 5.0 GHz CLOCK SPEED
Core 0ALU + FPU
Core 1ALU + FPU
Core 2ALU + FPU
Core 3ALU + FPU
MASSIVE L3 CACHE (~256 MB) + ADVANCED BRANCH PREDICTOR
DDR5 SYSTEM BUS β€” OPTIMIZED FOR SERIAL LATENCY
The Metaphor: An Einstein-level mind that solves one hyper-complex problem at a time.
THE PARALLEL BRUTE-FORCE ENGINE

Meet the GPU: 5,000 Line Cooks

Now replace that single chef with an army of 5,000 junior line cooks in a stadium kitchen. Can they design an intercontinental dish? No. But every single one can flip a burger!

πŸ”
Massive Throughput: Need to flip 50,000 burgers? All 5,000 line cooks flip simultaneously. The order is done in 10 seconds.
πŸš€
Thousands of Cores: Modern GPUs (like H100) feature 14,000+ CUDA cores and hundreds of specialized Tensor Cores.
🌊
Terabytes/sec Memory Highway: High Bandwidth Memory (HBM) delivers up to 3.35 Terabytes/sec β€” a firehose of numbers feeding all cores.
GRAPHICS PROCESSING UNIT DIE (14,000+ CORES) 1.8 GHz β€’ SIMD ARCHITECTURE
SM #1
SM #2
SM #3
SM #4
SM #5
SM #6
SM #7
SM #8
HBM3e (80GB)
3,350 GB/s MEMORY BANDWIDTH
HBM3e (80GB)
The Metaphor: An army of thousands doing simple arithmetic in unison.
LIVE INTERACTIVE BENCHMARK

The 4K Pixel Race: Serial vs. Parallel

Task: Calculate lighting color for 600 independent pixels. Watch how CPU sequential execution compares to GPU parallel execution.

CPU: Serial Execution (1 Core)
Time: 0.00 ms
Clock: 5.0 GHz Active Workers: 1 Core Status: Ready
GPU was --x faster!
GPU: Parallel Execution (Massive Cores)
Time: 0.00 ms
Clock: 1.8 GHz Active Workers: 600 Cores Status: Ready
πŸ’‘

Why did the GPU destroy the CPU? Because pixel #47 has zero dependency on pixel #48! Neither core needs to wait for the other. This is called an Embarrassingly Parallel Workload.

THE BIG REVEAL

AI Isn't Magic. It's High School Linear Algebra!

When ChatGPT predicts the next token, it isn't contemplating existence. It is computing billions of simple multiplications and additions: y = W Β· x + b.

1. Giant Model Weights [W]
w₁₁
w₁₂
w₁₃
w₁₄
w₂₁
wβ‚‚β‚‚
w₂₃
wβ‚‚β‚„
w₃₁
w₃₂
w₃₃
w₃₄
w₄₁
wβ‚„β‚‚
w₄₃
wβ‚„β‚„
70 Billion Learned Parameters in Llama-3
βœ•
2. Input Vector [x]
x₁
xβ‚‚
x₃
xβ‚„
User prompt token embeddings
=
3. Output Activation [y]
y₁ = (w₁₁·x₁) + (w₁₂·xβ‚‚) + (w₁₃·x₃) + (w₁₄·xβ‚„) + b₁
Multiply & Accumulate (MAC)

No decision trees. No complex if/else. Just millions of independent dot products simultaneously!

ENGINEERING ARCHITECTURE

CPU vs. GPU: Architectural Showdown

Neither is "better" β€” they are engineered for fundamentally opposite computing goals.

Architectural Feature Central Processing Unit (CPU) Graphics Processing Unit (GPU) Why It Matters for AI
Core Architecture 8 to 64 large, powerful cores 14,000+ tiny parallel cores AI matrices split across thousands of parallel workers.
Optimization Target Latency (Finish 1 task ASAP) Throughput (Finish 1,000,000 tasks) LLMs don't need 1 fast word; they need billions of weights computed together.
Memory Bandwidth ~100 to 300 GB/s (DDR5) Up to 3,350 GB/s (HBM3e) Prevents compute cores from starving waiting for memory weights.
Branching & Control Logic Very high (Branch Predictors, Out-of-Order) Minimal (Rigid SIMD / Lockstep) Neural networks follow deterministic matrix pipelines without branching.
Role in AI Infrastructure The Orchestrator: OS, networking, tokenizers, data loaders The Math Engine: Matrix multiplies, backpropagation, KV cache Modern AI servers pair host CPUs with 8x GPUs via PCIe/NVLink.
CHAPTER 01 SUMMARY

Three Golden Rules to Remember

1

CPUs are Master Chefs

Built for low latency, complex branching, OS kernels, and sequential execution.

2

GPUs are 5,000 Line Cooks

Built for massive parallel throughput, high memory bandwidth, and simple repeated math.

3

AI Math is Easy to Do in Parallel

AI involves billions of matrix multiplications that can be done at the same time. That’s exactly what GPUs are built for

UP NEXT: CHAPTER 02

Inside a GPU: CUDA Cores, Tensor Cores & VRAM

What actually happens inside an NVIDIA chip during matrix math? Why did Tensor Cores 10x AI speeds? And why does GPU VRAM cost so much?

CUDA Core
Tensor Core (FP8 / FP16)
HBM VRAM Pool