Why Did AI Fall in Love With GPUs?
When OpenAI built the clusters to train GPT-4, they didn't buy thousands of 64-core Intel or AMD server CPUs. They spent hundreds of millions of dollars buying tens of thousands of GPUs.
64-Core Server CPU
The undisputed king of data centers for 30 years. Runs Linux, databases, and microservices.
NVIDIA H100 Tensor Core GPU
Originally engineered for rendering 3D video games like Cyberpunk 2077.
Meet the CPU: The Michelin-Star Chef
Imagine the world's most brilliant Master Chef. They can invent complex recipes, taste and adjust delicate seasoning, manage kitchen inventory, and talk to restaurant guests simultaneously.
if/else statements, context switches, and interrupts without breaking a sweat.
Meet the GPU: 5,000 Line Cooks
Now replace that single chef with an army of 5,000 junior line cooks in a stadium kitchen. Can they design an intercontinental dish? No. But every single one can flip a burger!
The 4K Pixel Race: Serial vs. Parallel
Task: Calculate lighting color for 600 independent pixels. Watch how CPU sequential execution compares to GPU parallel execution.
Why did the GPU destroy the CPU? Because pixel #47 has zero dependency on pixel #48! Neither core needs to wait for the other. This is called an Embarrassingly Parallel Workload.
AI Isn't Magic. It's High School Linear Algebra!
When ChatGPT predicts the next token, it isn't contemplating existence. It is computing billions of simple multiplications and additions: y = W Β· x + b.
No decision trees. No complex if/else. Just millions of independent dot products simultaneously!
CPU vs. GPU: Architectural Showdown
Neither is "better" β they are engineered for fundamentally opposite computing goals.
| Architectural Feature | Central Processing Unit (CPU) | Graphics Processing Unit (GPU) | Why It Matters for AI |
|---|---|---|---|
| Core Architecture | 8 to 64 large, powerful cores | 14,000+ tiny parallel cores | AI matrices split across thousands of parallel workers. |
| Optimization Target | Latency (Finish 1 task ASAP) | Throughput (Finish 1,000,000 tasks) | LLMs don't need 1 fast word; they need billions of weights computed together. |
| Memory Bandwidth | ~100 to 300 GB/s (DDR5) | Up to 3,350 GB/s (HBM3e) | Prevents compute cores from starving waiting for memory weights. |
| Branching & Control Logic | Very high (Branch Predictors, Out-of-Order) | Minimal (Rigid SIMD / Lockstep) | Neural networks follow deterministic matrix pipelines without branching. |
| Role in AI Infrastructure | The Orchestrator: OS, networking, tokenizers, data loaders | The Math Engine: Matrix multiplies, backpropagation, KV cache | Modern AI servers pair host CPUs with 8x GPUs via PCIe/NVLink. |
Three Golden Rules to Remember
CPUs are Master Chefs
Built for low latency, complex branching, OS kernels, and sequential execution.
GPUs are 5,000 Line Cooks
Built for massive parallel throughput, high memory bandwidth, and simple repeated math.
AI Math is Easy to Do in Parallel
AI involves billions of matrix multiplications that can be done at the same time. Thatβs exactly what GPUs are built for