Conversation
AI-Assisted: GPT-5.6 Sol + Sentinel
AI-Assisted: GPT-5.6 Sol + Sentinel
AI-Assisted: GPT-5.6 Sol + Sentinel
Restore the pre-rewrite cache, measurement, W4, SIMD, fusion, and profiling teaching while retaining the current executable checkpoints and accepted evidence. AI-Assisted: GPT-6 Sol + Sentinel
Correct the Week 2 cache, SIMD, and fused-kernel measurement steps
Map learner and reference tests, checkpoint labels, extension TODOs, and book commands to the five-day progression while preserving nine checkpoint routes and existing implementations. AI-Assisted: GPT-6 Sol + Forge
Replace the rejected abstract progression label with the concrete selected-engine run and align its exact helper assertion. AI-Assisted: GPT-6 Sol + Forge
Fix Week 2 entry prerequisites and retired decode address
Use independent MLX attention for the focused tiled-prefill expectation and align cooperative tile TODO ownership with the Day 3 lesson. AI-Assisted: GPT-6 Sol + Forge
Name retired bounded decode in Week 2 sidebar
This was referenced Sep 24, 2026
skyzh
added a commit
that referenced
this pull request
Sep 24, 2026
## Scope Extract Week 2 Day 2 from the reviewed five-day Draft #316 onto the merged Day 1 course. The active lesson teaches packed W4 weights and the `quantized-matvec` checkpoint after the Day 1 capacity cache. Starter, reference code, supplied tests, CLI, benchmarks, profiler, and shipped-day CI expose the Day 2 contract while reserving later kernels and checkpoint names for later PRs. The lesson retains the signed-affine W4 derivation, eight-nibble packing, BF16 metadata, Apple bandwidth and roofline example, matched benchmark method, and historical Week 2 URLs. Readable text accompanies the displayed equations. The preface identifies Day 1 as the dense control and Day 2 as the packed path; the 4B comparison is optional. ## Local evidence - Both learner and reference native extensions build. The Day 2 shipped-day gate collected 351 tests: 350 passed, 1 skipped. Day 1 and Day 2 reference files passed 6/6 and 13/13. Copied learner tests collect and stop at named TODO seams. - Installation, documented model-free command parsing, mdBook, sitemap, rendered local links, lint, diff and Git integrity checks passed on the integrated local head. The two-page documentation correction was checked again on its direct successor. - Independent local reviews on the corrected head `fabba526` returned factual/preservation PASS, math/command correctness GO, learner WALKABLE, and static accessibility PASS. These verdicts are local evidence; this public PR still requires fresh exact-head reviews. ## Open gates Hosted macOS run 35950765030 succeeded on exact head `fabba526`: both native extensions built and the shipped-day manifest collected 351 tests (349 passed, 2 skipped). The runner's skip count differs from the local author run's one skip. This PR is Draft. Fresh public-head factual, correctness, learner, and accessibility reviews and all current repository gates remain required before Ready or merge. No Qwen model inference or performance benchmark was run for this Day 2 slice, and this PR makes no speed claim. The separate full-course Draft #316 performance screen does not validate Day 2 performance. The current main-branch workflow deploys GitHub Pages on push. Merging this PR would trigger publication, so the Day 2 publication decision remains separate from the approval used for Day 1. No release, tag, settings, visibility, or later-day work is part of this PR. --------- Co-authored-by: Sentinel <sentinel@raft.local> Co-authored-by: Forge <forge@raft.local>
skyzh
added a commit
that referenced
this pull request
Sep 24, 2026
## Scope Add Week 2 Day 3 SIMD matrix prefill on top of the merged Day 2 W4 course. The active lesson introduces the `simd-matmul` checkpoint, a 10×97 partial-tile operator exercise, and a public model-output comparison after the completed Day 1–2 checkpoints. The learner-owned native SIMD seam remains unfinished in the starter; reference code and supplied tests provide the runnable control. CLI, benchmark selectors, and shipped-day CI expose Day 3 while later-day routes remain deferred. The chapter explains the 32×32 cooperative tile, four 16×16 quadrants, guarded edge stores, readable equation alternative, and matched measurement method. Former Week 2 URLs remain historical. Required measurement uses the setup-cached 0.6B model; 4B is optional. This PR makes no speed claim. ## Local evidence - Both learner and reference extensions build; installation checks pass. The Day 3 shipped-day gate collected 369 tests: 368 passed, 1 optional-model skip. Direct Day 1–3 reference files passed 25 tests. - Supplied Task 1 compares public 1×3 outputs on the fallback path. Task 2 checks the 10×97 SIMD operator against the readable packed control. Task 3 uses fixed tiny fixtures with seeds 0 and 4, compares public 1×10 checkpoint BF16 outputs and full shapes with separate caches at `atol=0.75`, `rtol=0.05`, and restores MLX RNG state. The unfinished learner starter reaches its named SIMD TODO in Task 2; cumulative model tasks still require Days 1–2 completion. - mdBook, sitemap, rendered links, chapter command parsing, lint, diff hygiene, and Git integrity checks passed on local head `49bda24c`. Fresh independent local reviews returned factual/preservation PASS, correctness/public-contract GO, learner WALKABLE, and static accessibility PASS. These are local verdicts only. ## Open gates Hosted macOS run 35969966892 succeeded on exact head `49bda24c`: both native extensions built and the shipped-day manifest collected 369 tests (367 passed, 2 skipped). The runner has one more skip than the local author run; independent CI review will interpret the difference. This PR is Draft. Fresh independent reviews of this exact public head, base, and body and all current repository gates remain required before Ready or merge. No cached Qwen benchmark was run for this Day 3 slice; the separate full-course Draft #316 performance screen does not establish Day 3 speed. Chi approved normal website updates for this five-day Week 2 batch. The ordinary main-branch Pages workflow may run after a separately gated merge. This PR does not authorize a release, tag, settings or visibility change, manual deployment, or later-day work. --------- Co-authored-by: Sentinel <sentinel@raft.local> Co-authored-by: Forge <forge@raft.local>
skyzh
added a commit
that referenced
this pull request
Sep 24, 2026
## Scope Add Week 2 Day 4 compact model primitives on top of merged Day 3. The active lesson, starter seams, readable reference controls, supplied tests, CLI selectors, benchmarks, and shipped-day manifest progress cumulatively through `rmsnorm`, `rope`, and `swiglu` after `simd-matmul`. Learners implement the register-cached RMSNorm kernel and its wide fallback, RoPE, and SwiGLU; Tasks 1–3 also assign the three native CPU evaluators to report GPU-only errors. Day 5 tiled attention and `selected` remain deferred. The chapter preserves the RMSNorm reduction stages, RoPE head-pair math, SwiGLU equation, cumulative model checks, and matched measurement method. Former Week 2 URLs not reused for active Day 4 remain historical. Required measurement uses the setup-cached Qwen3-0.6B model; a 4B repeat is optional. This PR makes no speed claim. ## Local evidence - The frozen tree is `a4e49d7de3c281f8ada41b5c04c0482283bcd7bc` at head `18613b83bb4e743f1bae89bde02f8950b4947166`, directly descended from merged Day 3 main `df51f93acf1ba24e3aa02d1460d870bd04dc7978`. The cumulative diff is 31 Day 4 paths. The final lesson correction changed only one page; engineering and tests retain their accepted combined blobs. - Both native extensions build. On the final head, independent review passed the Day 4 reference selector (7/7), the Day 4/RMSNorm boundary/interface set (41/41), and the unchanged extension interface test. Copied learner tests reach named Tasks 1–3 TODOs; unfinished cumulative model tasks remain expected learner failures. The shipped-day manifest collects 398 nodes with Day 5 deferred. - On the preceding combined author tree, the printed shipped-day gate passed 397 tests with one optional-model skip. The final one-page successor passed its focused, book, and integrity checks. Hosted macOS run `35979136982` succeeded on this exact public head: 398 tests collected, 396 passed, 2 skipped. The hosted log does not name the skipped nodes. - mdBook, sitemap, rendered local links, command syntax, diff hygiene, and strict Git integrity passed. Four fresh independent local reviews on the final head returned factual/preservation GO, correctness/test GO, learner WALKABLE, and static accessibility PASS. These are local verdicts only. ## Open gates This PR is Draft. Hosted macOS shipped-day CI has passed on this head. Fresh independent reviews of the exact public head, base, and body, interpretation of the hosted skips, and current repository rules remain gates before Ready or ordinary merge. No Day 4 performance benchmark was run; the separate full-course Draft #316 screen does not establish a Day 4 speedup. Chi approved normal website updates for this five-day Week 2 batch. The ordinary main-branch Pages workflow may run after a separately gated merge. This PR does not authorize a release, tag, settings or visibility change, manual deployment, or later-day work. --------- Co-authored-by: Sentinel <sentinel@raft.local>
skyzh
added a commit
that referenced
this pull request
Sep 24, 2026
## Scope Add Week 2 Day 5 tiled dense-attention prefill and the cumulative `selected` checkpoint on merged Day 4. Learners implement the BF16/D128 native tiled path for query length at least 9, with causal/additive masks and grouped-query attention. Short or unsupported cases use the readable fallback. The `selected` checkpoint exposes capacity, RMSNorm, and tiled-attention controls independently. Required matched measurement uses the setup-cached Qwen3-0.6B model; 4B is optional. This PR makes no speed claim. The lesson makes Days 1–5 active and qualifies older Week 2 URLs that were reused by active lessons. The repaired Day 1–2 supplied tests check public cache, quantized-operator, and packed-model behavior instead of private layout, source text, or cache movement counters; the Day 2 lesson describes that observable scope. ## Local evidence - Frozen head `360e485860b387ba7ecdc32bf24113cb269a7798`, tree `7b1159092980fa82ff8a453aea259ba829747c54`, descended from merged Day 4 main `68642f05df1094ca20f92779504f737921a8dbd9`. The cumulative diff is exactly 55 paths, each matched to an accepted source blob. The final integration commit changes only the two repaired supplied-test files on the corrected docs parent. - Both native extensions and installation passed. The full unfiltered reference suite passed 679 tests with 2 optional skips; direct Day 1–5 selectors passed 6/14/6/7/10. Focused Day 5, tiled-boundary, interface, and repaired Day 1–2 tests passed in independent review. Copied learner tests reach named native and preceding-day TODOs. - mdBook/sitemap rendered 49 pages; the final-tree local scan found 1,909 links and 993 fragments with no broken targets. The 39 lesson Bash blocks passed syntax checks, and nine focused selectors collected the intended tests. Changed-file lint/format, diff hygiene, and strict Git checks passed. - Four fresh independent local reviews on this exact head returned factual/preservation PASS, correctness/test GO, learner WALKABLE (conditional on completing Days 1–4), and static-accessibility PASS. These are local verdicts only. No product benchmark or real-model speed result is claimed. ## Open gates This PR is Draft. Hosted macOS full unfiltered reference CI succeeded on this exact head (681 collected, 677 passed, 4 optional model skips in Week 1 Day 5 and Week 3 Day 1); the required 0.6B tests ran. Four fresh independent reviews of the exact public head, base, and current body remain required before Ready or ordinary merge. The full-course Draft #316 remains source material; its verdicts do not transfer. Chi approved normal website updates for the five-day Week 2 batch. The ordinary main-branch Pages workflow may run after a separately gated merge. This PR does not authorize release, tag, settings or visibility change, manual deployment, or later-course work.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Exact candidate
61666ee1290edf49444f7a936ac7955240e630f3c4f2d5e4672a43c766e62abf8ba1108a8465427amainatd3b704fd6de51647cee48a2a85ae9fc3cb22df8cThe earlier seven-day preview on this PR,
7d02ffaa, is superseded. Its independent verdicts and hosted result do not transfer to this head. Draft PR #315 remains an older, divergent route and is untouched.Five-day route
selectedinference-engine run.The book was rebuilt from the pre-rewrite public draft. It retains the cache worked example, benchmark method, W4 math, Mac-chip bandwidth/roofline reasoning, SIMD/fusion mechanisms, and optional macOS profiling workflow. Retired decode and Split-K experiments are labeled historical. The learner owns the cache and kernels; MLX is a baseline or oracle.
The starter, reference, tests, CLI, benchmark/profiler/capture labels, and book commands now use the five-day route and nine named checkpoints. Historical benchmark records keep their original labels. Week 3/4 implementation is unchanged.
Local evidence and open gates
On the integrated ancestor
027cb3e8, the model-free reference suite passed 651 cases with 16 model-loading instances deselected; five focused reference-day gates passed 6/10/2/4/7; 114 helper cases, extension builds, mdBook/sitemap, changed-page links, lint, and learner TODO diagnostics were checked. Later author corrections repaired the historical decode URL, Week 1/Day 1 extension and focused-gate instructions, the Day 5 independent test oracle, and Day 3 native labels. On the current head, Forge reported five focused reference-day gates 6/10/2/4/7, 124 affected helpers, both extension builds, lint, mdBook/sitemap, and 43 corrected-page links. These author-run results do not stand in for independent review.On parent head
c0b3871d, Oracle’s factual/preservation review PASS (task #734), Sage’s correctness review GO (task #735), and Scholar’s learner review WALKABLE (task #736) were accepted; its hosted macOS reference check succeeded. Herald’s static accessibility review BLOCK (task #737) found thatbook/src/SUMMARY.mdmislabeled the retired bounded-decode page. The one-line correction in task #738 is now this PR head: “Retired: Bounded Decode Attention,” with its link target preserved. Those parent-head verdicts and the hosted result are historical. On this exact61666ee1head, Oracle factual/preservation PASS (task #739), Sage correctness GO (task #740), Scholar learner WALKABLE (task #741), Herald static accessibility PASS (task #742), and the hosted macOS reference check are accepted/successful. Browser and assistive-technology behavior remains untested; no prior-head verdict was transferred. The still-oldercef0adfcevidence is historical too.The canonical Qwen3-4B product-performance matrix is still missing a clean terminal result, including the required 2K/512 and 8K rows; no 80%-of-MLX claim is made. Task #722 used its one strict quiet-host admission on this frozen head in Chi’s approved Pacific window. Time Machine was copying in all five snapshots and Spotlight indexing exceeded the CPU noise limit in three, so the attempt is terminal INCONCLUSIVE with zero model work and no retry. The signed task #722 report (4,658 bytes, SHA-256
aba8f5f18e3be2f8904493add4ec1f91dc2d1a0128d514d2d8336013dcf22854) and raw snapshot/manifest integrity were independently checked and accepted. No further admission is authorized under that task. The prior task #709 admission also failed before model samples and cannot fill this gap.This Draft authorizes no Ready transition, merge, release, tag, deployment, publication, settings change, visibility change, or later-course work.