Why is Reading Fast, but Writing Slow?
You paste a 2,000-word document into an LLM. It ingests the entire prompt in 100 milliseconds. Yet generating a short 50-word response takes 5 to 10 seconds! Why?
The Speed-Reader
2,000 words swallowed at once. All tokens processed simultaneously in parallel.
The Typewriter
50 words generated one... by... one. Autoregressive feedback loop.
The Prefill Phase: The Speed-Reader
Because you provided all 2,000 words upfront, the GPU doesn't have to guess. It reads and multiplies the entire sequence simultaneously!
[Batch, SeqLen, Dim] Γ [Weights].
The Decode Phase: The Typewriter
Why can't the GPU generate 50 response words at once? Because word #2 requires knowing word #1. Generation is strictly sequential!
[1, Dim] Γ [Weights].
Interactive Simulator: Prefill vs. Decode
Watch an LLM execute in real-time. Experience the instantaneous burst of parallel Prefill followed by the memory-bound typewriter Decode stream.
Notice the Dramatic Shift: During Prefill, Compute hits 94% with peak efficiency. But during Decode, Compute collapses to 4% while Memory Bandwidth hits 88%!
Streaming 70GB for ONE Word!
To generate one single token from a 70-billion parameter model, the GPU must move all 70GB of weights from VRAM to the compute cores.
Generating 30 tokens/sec on Llama-3 70B requires:
70 GB Γ 30 tokens/sec = 2,100 GB/s (2.1 TB/s!)
How vLLM Beats the Decode Bottleneck
Modern inference engines use two genius infrastructure techniques to turn memory-bound waste into maximum throughput.
Continuous Batching
Running Decode for 1 user wastes 95% of your GPU compute. By batching 64 users together, that 70GB weight stream calculates 64 tokens in the same memory pass!
Speculative Decoding
A tiny, ultra-fast 1B "draft model" guesses 5 tokens ahead. The giant 70B model verifies all 5 guesses in 1 parallel Prefill pass!
Three Golden Takeaways
Prefill is Compute-Bound (GEMM)
Processes all prompt tokens in parallel at peak Tensor Core TFLOPS in milliseconds.
Decode is Memory-Bound (GEMV)
Generates one token per step. Must stream all model weights from VRAM for each word.
HBM & Batching Rule Generative AI
Higher memory bandwidth and continuous batching are what make LLM servers profitable.