Commit 8e332ed
committed
KV cache post: cut to ~8min, drop math/jargon, classroom analogy from top
Big simplification pass for non-ML readers:
- Cut: embeddings-as-vectors section, full attention formula, multi-head
matrix details, transformer block decomposition (LayerNorm/FFN/residual),
causal masking matrix diagram, naive cost big-O analysis, full memory
math breakdown by dimensions.
- Kept: tokenization (one paragraph), classroom + Q/K/V cards, layers and
parallel classrooms (brief), generation loop, KV cache (with SVG diagram),
memory cost (with calculator widget), modern serving optimizations as
folder operations, one-breath summary.
- Read time: 14 min → 8 min.1 parent e428265 commit 8e332ed
1 file changed
Lines changed: 81 additions & 186 deletions
0 commit comments