THE PROMPTING PARADOX

Why is Reading Fast, but Writing Slow?

You paste a 2,000-word document into an LLM. It ingests the entire prompt in 100 milliseconds. Yet generating a short 50-word response takes 5 to 10 seconds! Why?

PHASE 1: PROMPT INGESTION ⚑ 120 ms
πŸ“–

The Speed-Reader

2,000 words swallowed at once. All tokens processed simultaneously in parallel.

VS
PHASE 2: TOKEN GENERATION 🐒 6,200 ms
⌨️

The Typewriter

50 words generated one... by... one. Autoregressive feedback loop.

πŸ€” Let's look inside the silicon to see why this happens on every GPU in the world.
PHASE 1: PARALLEL COMPUTATION

The Prefill Phase: The Speed-Reader

Because you provided all 2,000 words upfront, the GPU doesn't have to guess. It reads and multiplies the entire sequence simultaneously!

πŸ“¦
Matrix-Matrix Multiply (GEMM): All 2,000 tokens are packed into a 2D matrix: [Batch, SeqLen, Dim] Γ— [Weights].
⚑
Compute-Bound (Peak TFLOPS): Arithmetic intensity is extremely high (>100 FLOPs/byte). Tensor Cores are saturated at 95%+ utilization!
πŸ“
Populates the KV Cache: It calculates attention for all prompt tokens once and stores Key & Value vectors in VRAM for Phase 2.
PREFILL EXECUTION PATTERN MATRIX Γ— MATRIX
Input Tokens Matrix (2,000 tokens)
tok 1tok 2tok 3...
tok 500tok 501tok 502...
tok 1997tok 1998tok 1999tok 2000
βœ• [Weights Matrix]
πŸ”₯ ALL TENSOR CORES FIRING SIMULTANEOUSLY
Result: Swallows 2,000 tokens in under 100 milliseconds.
PHASE 2: AUTOREGRESSIVE LOOP

The Decode Phase: The Typewriter

Why can't the GPU generate 50 response words at once? Because word #2 requires knowing word #1. Generation is strictly sequential!

πŸ”„
Autoregressive Feedback Loop: Every step produces exactly 1 single token, which loops back into the input for the next step.
πŸ“‰
Matrix-Vector Multiply (GEMV): Input is no longer a big matrixβ€”it is a tiny 1D vector: [1, Dim] Γ— [Weights].
🚧
The Memory Bandwidth Trap: Arithmetic intensity plummets to ~1 FLOP/byte. Tensor Cores drop to 4% utilization waiting for memory!
DECODE STEP PIPELINE MATRIX Γ— 1 VECTOR
Token N
βž” Forward Pass βž”
Token N+1
⚠️ TENSOR CORES IDLE WAITING FOR VRAM WEIGHTS
Result: Speed is capped by memory bandwidth, not compute TFLOPS.
LIVE EXECUTION SIMULATOR

Interactive Simulator: Prefill vs. Decode

Watch an LLM execute in real-time. Experience the instantaneous burst of parallel Prefill followed by the memory-bound typewriter Decode stream.

Status: Ready to execute
Inference Stream 0.0 tok/s
[Prompt: 500 tokens loaded] β–Œ
SILICON HARDWARE TELEMETRY
Tensor Core Compute: 0%
HBM Memory Bandwidth: 0%
Phase Mode: Idle
Arithmetic Intensity: -- FLOPs/byte
πŸ’‘

Notice the Dramatic Shift: During Prefill, Compute hits 94% with peak efficiency. But during Decode, Compute collapses to 4% while Memory Bandwidth hits 88%!

THE SILICON REALITY

Streaming 70GB for ONE Word!

To generate one single token from a 70-billion parameter model, the GPU must move all 70GB of weights from VRAM to the compute cores.

🀯
The Math of Autoregression:
Generating 30 tokens/sec on Llama-3 70B requires:
70 GB Γ— 30 tokens/sec = 2,100 GB/s (2.1 TB/s!)
🌊
Niagara Falls on Silicon: 2.1 Terabytes of numbers must rush through the memory bus every single second just to type text!
βš–οΈ
Why HBM Rules: An RTX 4090 caps out at 1,000 GB/s (cannot do 30 tok/s on 70B). An H100 delivers 3,350 GB/s, enabling fluid real-time speed.
70B MODEL DECODE BANDWIDTH DRAIN 2.1 TB/s CONSUMED
70GB MODEL WEIGHTS IN VRAM
2,100 GB/s CONTINUOUS MEMORY STREAM
14,000 CORES β†’ PRODUCES 1 TOKEN
Takeaway: Token generation speed is governed by memory speed, not compute speed!
SERVING ARCHITECTURE

How vLLM Beats the Decode Bottleneck

Modern inference engines use two genius infrastructure techniques to turn memory-bound waste into maximum throughput.

TECHNIQUE 1

Continuous Batching

Running Decode for 1 user wastes 95% of your GPU compute. By batching 64 users together, that 70GB weight stream calculates 64 tokens in the same memory pass!

Up to 10x Server Throughput
TECHNIQUE 2

Speculative Decoding

A tiny, ultra-fast 1B "draft model" guesses 5 tokens ahead. The giant 70B model verifies all 5 guesses in 1 parallel Prefill pass!

2.5x Faster Latency (Zero Quality Loss)
CHAPTER 04 SUMMARY

Three Golden Takeaways

1

Prefill is Compute-Bound (GEMM)

Processes all prompt tokens in parallel at peak Tensor Core TFLOPS in milliseconds.

2

Decode is Memory-Bound (GEMV)

Generates one token per step. Must stream all model weights from VRAM for each word.

3

HBM & Batching Rule Generative AI

Higher memory bandwidth and continuous batching are what make LLM servers profitable.

UP NEXT: CHAPTER 05

GPU vs. TPU vs. CPU: Which One Should You Use?

Is NVIDIA the only game in town? What are Google TPUs with Systolic Arrays? And when does running models on cheap CPUs actually save millions of dollars?

NVIDIA GPU (CUDA Ecosystem) Google TPU (Systolic Arrays) Intel/AMD CPU (Cost Optimizer)