Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 2 additions & 4 deletions .github/workflows/macos.yml
Original file line number Diff line number Diff line change
Expand Up @@ -38,12 +38,10 @@ jobs:
- name: Try building extensions
run: |
pdm run build-ext
pdm run build-ext-test

- run: pdm run build-ext-ref

- run: cargo install mdbook-toc
- run: cargo install mdbook-katex --version 0.10.0-alpha
- uses: taiki-e/install-action@mdbook

- run: pdm run test-refsol
- name: Run printed shipped-day reference CI manifest
run: pdm run python scripts/week2_shipped_day_ci.py
20 changes: 7 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,12 +27,10 @@ The course follows a four-week learning path:
- **Week 1: From Matmul to Text.** Build a Qwen3 model directly from `mlx.core`
array operations: attention, RoPE, GQA, RMSNorm, the MLP, sampling, and
the autoregressive loop.
- **Week 2: A Step Closer to vLLM.** Add a KV cache, establish a
synchronized MLX baseline, and let matched benchmarks choose each
optimization. The causal path moves from quantized decode matvec to fused
model kernels and SIMD-matrix prefill; decode attention is an optional
workload-conditioned lab, and split-K stays only where a measured short
shape supports it.
- **Week 2: A Faster Single Request.** The current Day 1 route adds
`kv-cache` and request-bounded `capacity-cache`, then measures both against
the Week 1 full-prefix control. Later packed-W4, SIMD, fused-primitive, and
tiled-attention lessons are planned; their checkpoints are not Day 1 gates.
- **Week 3: Build a Mini vLLM.** Introduce continuous
batching and chunked admission, then make paged KV the canonical serving
layout. Decode attention and FlashAttention learn to read pages directly so
Expand Down Expand Up @@ -109,13 +107,7 @@ one explicit byte range through the existing loop.
| 1.5 | Load the Model | ✅ | ✅ | ✅ | ✅ |
| 1.6 | Generate Responses (aka Decoding) | ✅ | ✅ | ✅ | ✅ |
| 1.7 | Sampling | ✅ | ✅ | ✅ | ✅ |
| 2.1 | KV Cache | ✅ | ✅ | ✅ | 🚧 |
| 2.2 | Benchmarking and Profiling | ✅ | ✅ | ✅ | 🚧 |
| 2.3 | Quantize the Model | ✅ | ✅ | ✅ | 🚧 |
| 2.4 | Fused Model Kernels | ✅ | ✅ | ✅ | 🚧 |
| 2.5 | SIMD-Matrix Prefill | ✅ | ✅ | ✅ | 🚧 |
| 2.6 (optional) | Workload-Conditioned Operator Lab | ✅ | ✅ | ✅ | 🚧 |
| 2.7 | Conditional Split-K and Final Decision | ✅ | ✅ | ✅ | 🚧 |
| 2.1 | Cache and Measure (`kv-cache`, `capacity-cache`) | 🚧 | 🚧 | ✅ | 🚧 |
| 3.1 | Continuous Batching | ✅ | ✅ | ✅ | 🚧 |
| 3.2 | Chunked Prefill | ✅ | ✅ | ✅ | 🚧 |
| 3.3 | Paged KV Cache | ✅ | ✅ | ✅ | 🚧 |
Expand All @@ -133,6 +125,8 @@ one explicit byte range through the existing loop.
| 4.8 | Fork, Steer, and Select | ✅ | ✅ | ✅ | 🚧 |
| 4.9 | Bound Tool Evidence | ✅ | ✅ | ✅ | 🚧 |

The older Week 2 chapter URLs remain available as [historical material](book/src/week2-02-benchmark-profile.md). Their former Day 2–7 tests and checkpoint commands are not part of the current Day 1 learner route.

Other topics not covered include quantized or compressed KV caches,
cross-request prefix caching, fine-tuning, and long-context techniques.

Expand Down
20 changes: 7 additions & 13 deletions benches/bench.py
Original file line number Diff line number Diff line change
Expand Up @@ -86,16 +86,7 @@ def parse_args() -> argparse.Namespace:
)
parser.add_argument(
"--week2-checkpoint",
choices=(
"kv-cache",
"quantized-matvec",
"rmsnorm",
"rope",
"swiglu",
"simd-matmul",
"decode-attention",
"split-k",
),
choices=("kv-cache", "capacity-cache"),
help="run one cumulative Week 2 end-to-end checkpoint",
)
parser.add_argument("--device", type=str, default="gpu", choices=["cpu", "gpu"])
Expand Down Expand Up @@ -165,12 +156,15 @@ def validate_args(args: argparse.Namespace) -> None:
and args.device != "gpu"
and (
args.loader == "week3"
or (args.loader == "week2" and args.week2_checkpoint != "kv-cache")
or (
args.loader == "week2"
and args.week2_checkpoint not in (None, "kv-cache", "capacity-cache")
)
)
):
raise ValueError(
"The completed Week 2 and Week 3 custom-kernel models are GPU-only; "
"use the Week 2 kv-cache checkpoint for the readable pre-kernel path"
"Week 3 custom-kernel models are GPU-only; "
"Day 1 Week 2 checkpoints use the readable path"
)
if args.disable_paged_attention and args.loader != "week3":
raise ValueError("--disable-paged-attention requires --loader week3")
Expand Down
50 changes: 4 additions & 46 deletions benches/bench_course_progression.py
Original file line number Diff line number Diff line change
Expand Up @@ -42,59 +42,17 @@ class Throughput:
WEEK1_VARIANT,
Variant(
"week2-kv-cache",
"2.1 KV cache",
"2.1 Reuse the prefix",
"ref",
"week2",
("--week2-checkpoint", "kv-cache"),
),
Variant(
"week2-quantized-matvec",
"2.3 Quantized matvec",
"week2-capacity-cache",
"2.1 + Bound KV-cache movement",
"ref",
"week2",
("--week2-checkpoint", "quantized-matvec"),
),
Variant(
"week2-rmsnorm",
"2.4 Fast RMSNorm",
"ref",
"week2",
("--week2-checkpoint", "rmsnorm"),
),
Variant(
"week2-rope",
"2.4 + Fast RoPE",
"ref",
"week2",
("--week2-checkpoint", "rope"),
),
Variant(
"week2-swiglu",
"2.4 + Fused SwiGLU",
"ref",
"week2",
("--week2-checkpoint", "swiglu"),
),
Variant(
"week2-simd-matmul",
"2.5 SIMD matrix prefill",
"ref",
"week2",
("--week2-checkpoint", "simd-matmul"),
),
Variant(
"week2-decode-attention",
"2.6 Optional decode attention",
"ref",
"week2",
("--week2-checkpoint", "decode-attention"),
),
Variant(
"week2-split-k",
"2.7 Split-K prefill",
"ref",
"week2",
("--week2-checkpoint", "split-k"),
("--week2-checkpoint", "capacity-cache"),
),
MLX_VARIANT,
)
Expand Down
8 changes: 1 addition & 7 deletions benches/profile_week2_kernels.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,13 +25,7 @@

DEFAULT_CASES = (
"kv-cache:decode:128",
"quantized-matvec:decode:128",
"swiglu:decode:128",
"simd-matmul:prefill:128",
"simd-matmul:prefill:32",
"decode-attention:decode:128",
"decode-attention:prefill:128",
"split-k:prefill:32",
"capacity-cache:decode:128",
)
PROMPT_RULE = "synthetic-token-ids"
PREFILL_LOGITS = "all"
Expand Down
8 changes: 1 addition & 7 deletions benches/week2_gpudebug.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,13 +19,7 @@
PREFILL_LOGITS = "all"
KNOWN_CHECKPOINTS = (
"kv-cache",
"quantized-matvec",
"rmsnorm",
"rope",
"swiglu",
"simd-matmul",
"decode-attention",
"split-k",
"capacity-cache",
)


Expand Down
18 changes: 9 additions & 9 deletions book/src/SUMMARY.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,15 +13,15 @@
- [The Qwen3 Model](./week1-05-qwen3-model.md)
- [Generating the Response](./week1-06-generate-response.md)
- [Sampling and Preparing for Week 2](./week1-07-sampling-prepare.md)
- [🚧 Week 2: A Step Closer to vLLM](./week2-overview.md)
- [🚧 KV Cache](./week2-01-kv-cache.md)
- [🚧 Benchmarking and Profiling](./week2-02-benchmark-profile.md)
- [🚧 Optional: Metal Profiling](./week2-advanced-profiling.md)
- [🚧 Quantize the Model](./week2-03-quantize-model.md)
- [🚧 Fused Model Kernels](./week2-04-fused-model-kernels.md)
- [🚧 SIMD-Matrix Prefill](./week2-05-simd-matrix-prefill.md)
- [🚧 Day 6 (Optional): Workload-Conditioned Operator Lab](./week2-06-operator-lab.md)
- [🚧 Conditional Split-K and Final Decision](./week2-07-split-k-prefill.md)
- [🚧 Week 2: A Faster Single Request](./week2-overview.md)
- [🚧 Day 1: Cache and Measure](./week2-01-kv-cache.md)
- [Historical Week 2 lesson addresses](./week2-02-benchmark-profile.md)
- [Earlier quantization lesson](./week2-03-quantize-model.md)
- [Earlier fused-kernel lesson](./week2-04-fused-model-kernels.md)
- [Earlier SIMD-prefill lesson](./week2-05-simd-matrix-prefill.md)
- [Earlier bounded-decode lab](./week2-06-operator-lab.md)
- [Earlier Split-K lab](./week2-07-split-k-prefill.md)
- [Earlier optional macOS capture lab](./week2-advanced-profiling.md)
- [🚧 Week 3: Build a Mini vLLM](./week3-overview.md)
- [🚧 Continuous Batching](./week3-01-continuous-batching.md)
- [🚧 Chunked Prefill](./week3-02-chunked-prefill.md)
Expand Down
8 changes: 5 additions & 3 deletions book/src/appendix-performance.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,10 @@
# 🚧 Appendix: Performance Evidence Ledger

> **Status: Experimental, single-machine evidence.** See the
> [Week 2 verification matrix](./week2-overview.md#verification-status) before
> treating a correctness, integration, or performance result as broader proof.
> **Historical evidence from an earlier full Week 2 course state.** The
> [current Week 2 route](./week2-overview.md) ships Day 1 only. Later
> checkpoint labels, commands, and measured results below describe the older
> source tree; they are not runnable gates or performance results for this
> checkout.

This appendix records the measurements that determined the course order. The
numbers are not additive promises: after one bottleneck shrinks, every other
Expand Down
8 changes: 4 additions & 4 deletions book/src/course-roadmap.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
14 changes: 7 additions & 7 deletions book/src/glossary.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,13 +14,13 @@
- [Qwen3 Transformer Block](./week1-05-qwen3-model.md)
- [Week 1 Qwen3 Model](./week1-05-qwen3-model.md)
- [dequantize_linear](./week1-05-qwen3-model.md)
- [KV Cache](./week2-01-kv-cache.md)
- [Benchmarking and Profiling](./week2-02-benchmark-profile.md)
- [Quantize the Model](./week2-03-quantize-model.md)
- [Fused Model Kernels](./week2-04-fused-model-kernels.md)
- [Fused Decode Attention](./week2-05-decode-attention.md)
- [SIMD-Matrix Prefill](./week2-06-simd-matrix-prefill.md)
- [Split-K Prefill](./week2-07-split-k-prefill.md)
- [KV Cache and Request-Bounded Capacity](./week2-01-kv-cache.md)
- [Benchmarking, Profiling, and Decode Roofline](./week2-01-kv-cache.md#benchmark-the-cached-model)
- [Historical: Packed W4 Quantization](./week2-03-quantize-model.md)
- [Historical: Fused Model Kernels](./week2-04-fused-model-kernels.md)
- [Historical: SIMD-Matrix Prefill](./week2-05-simd-matrix-prefill.md)
- [Historical: Bounded Decode Attention](./week2-06-operator-lab.md)
- [Historical: Split-K Prefill](./week2-07-split-k-prefill.md)
- [Flash Attention](./week3-05-flash-attention.md)
- [Paged Attention](./week3-04-paged-attention-part2.md)

Expand Down
Loading