Skip to content

Commit 8e332ed

Browse files
committed
KV cache post: cut to ~8min, drop math/jargon, classroom analogy from top
Big simplification pass for non-ML readers: - Cut: embeddings-as-vectors section, full attention formula, multi-head matrix details, transformer block decomposition (LayerNorm/FFN/residual), causal masking matrix diagram, naive cost big-O analysis, full memory math breakdown by dimensions. - Kept: tokenization (one paragraph), classroom + Q/K/V cards, layers and parallel classrooms (brief), generation loop, KV cache (with SVG diagram), memory cost (with calculator widget), modern serving optimizations as folder operations, one-breath summary. - Read time: 14 min → 8 min.
1 parent e428265 commit 8e332ed

1 file changed

Lines changed: 81 additions & 186 deletions

File tree

0 commit comments

Comments
 (0)