Skip to content

Draft: Week 2 five-day integrated course (review in progress) - #316

Draft
skyzh wants to merge 22 commits into
mainfrom
atlas/week2-reviewed-preview-20260922
Draft

skyzh wants to merge 22 commits into
mainfrom
atlas/week2-reviewed-preview-20260922

Conversation

@skyzh

@skyzh skyzh commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Draft held for canonical Qwen3-4B performance evidence. This exact five-day head has fresh factual, correctness, learner, and static accessibility verdicts plus a successful hosted macOS check. The one authorized performance attempt failed its required quiet-host admission before model work. Keep this PR Draft; the 2K/512 and 8K product rows remain unavailable. Day-scoped extraction awaits a complete candidate decision.

Exact candidate

  • Head: 61666ee1290edf49444f7a936ac7955240e630f3
  • Tree: c4f2d5e4672a43c766e62abf8ba1108a8465427a
  • Base: public main at d3b704fd6de51647cee48a2a85ae9fc3cb22df8c
  • Diff from base: 66 files, +3,966/−3,502

The earlier seven-day preview on this PR, 7d02ffaa, is superseded. Its independent verdicts and hosted result do not transfer to this head. Draft PR #315 remains an older, divergent route and is untouched.

Five-day route

  1. Request-owned KV reuse and bounded append capacity, with matched Week 1 versus cached measurements.
  2. Packed W4 decode, including the quantization derivation, Apple memory-bandwidth table, and decode roofline.
  3. SIMD matrix prefill, with a control measured before the learner builds the candidate.
  4. RMSNorm, RoPE, and SwiGLU kernels with separate focused gates and cumulative model measurements.
  5. Tiled dense prefill attention, ending with a concrete selected inference-engine run.

The book was rebuilt from the pre-rewrite public draft. It retains the cache worked example, benchmark method, W4 math, Mac-chip bandwidth/roofline reasoning, SIMD/fusion mechanisms, and optional macOS profiling workflow. Retired decode and Split-K experiments are labeled historical. The learner owns the cache and kernels; MLX is a baseline or oracle.

The starter, reference, tests, CLI, benchmark/profiler/capture labels, and book commands now use the five-day route and nine named checkpoints. Historical benchmark records keep their original labels. Week 3/4 implementation is unchanged.

Local evidence and open gates

On the integrated ancestor 027cb3e8, the model-free reference suite passed 651 cases with 16 model-loading instances deselected; five focused reference-day gates passed 6/10/2/4/7; 114 helper cases, extension builds, mdBook/sitemap, changed-page links, lint, and learner TODO diagnostics were checked. Later author corrections repaired the historical decode URL, Week 1/Day 1 extension and focused-gate instructions, the Day 5 independent test oracle, and Day 3 native labels. On the current head, Forge reported five focused reference-day gates 6/10/2/4/7, 124 affected helpers, both extension builds, lint, mdBook/sitemap, and 43 corrected-page links. These author-run results do not stand in for independent review.

On parent head c0b3871d, Oracle’s factual/preservation review PASS (task #734), Sage’s correctness review GO (task #735), and Scholar’s learner review WALKABLE (task #736) were accepted; its hosted macOS reference check succeeded. Herald’s static accessibility review BLOCK (task #737) found that book/src/SUMMARY.md mislabeled the retired bounded-decode page. The one-line correction in task #738 is now this PR head: “Retired: Bounded Decode Attention,” with its link target preserved. Those parent-head verdicts and the hosted result are historical. On this exact 61666ee1 head, Oracle factual/preservation PASS (task #739), Sage correctness GO (task #740), Scholar learner WALKABLE (task #741), Herald static accessibility PASS (task #742), and the hosted macOS reference check are accepted/successful. Browser and assistive-technology behavior remains untested; no prior-head verdict was transferred. The still-older cef0adfc evidence is historical too.

The canonical Qwen3-4B product-performance matrix is still missing a clean terminal result, including the required 2K/512 and 8K rows; no 80%-of-MLX claim is made. Task #722 used its one strict quiet-host admission on this frozen head in Chi’s approved Pacific window. Time Machine was copying in all five snapshots and Spotlight indexing exceeded the CPU noise limit in three, so the attempt is terminal INCONCLUSIVE with zero model work and no retry. The signed task #722 report (4,658 bytes, SHA-256 aba8f5f18e3be2f8904493add4ec1f91dc2d1a0128d514d2d8336013dcf22854) and raw snapshot/manifest integrity were independently checked and accepted. No further admission is authorized under that task. The prior task #709 admission also failed before model samples and cannot fill this gap.

This Draft authorizes no Ready transition, merge, release, tag, deployment, publication, settings change, visibility change, or later-course work.

@skyzh skyzh changed the title Review preview: Week 2 capacity cache and tiled prefill course Draft: Week 2 integrated course (five-day book rewrite in progress) Sep 23, 2026
Sentinel and others added 4 commits September 23, 2026 00:28
Restore the pre-rewrite cache, measurement, W4, SIMD, fusion, and profiling teaching while retaining the current executable checkpoints and accepted evidence.

AI-Assisted: GPT-6 Sol + Sentinel
Correct the Week 2 cache, SIMD, and fused-kernel measurement steps
Map learner and reference tests, checkpoint labels, extension TODOs, and book commands to the five-day progression while preserving nine checkpoint routes and existing implementations.

AI-Assisted: GPT-6 Sol + Forge
Replace the rejected abstract progression label with the concrete selected-engine run and align its exact helper assertion.

AI-Assisted: GPT-6 Sol + Forge
@skyzh skyzh changed the title Draft: Week 2 integrated course (five-day book rewrite in progress) Draft: Week 2 five-day integrated course (review in progress) Sep 23, 2026
Sentinel and others added 3 commits September 23, 2026 01:38
Fix Week 2 entry prerequisites and retired decode address
Use independent MLX attention for the focused tiled-prefill expectation and align cooperative tile TODO ownership with the Day 3 lesson.

AI-Assisted: GPT-6 Sol + Forge
Name retired bounded decode in Week 2 sidebar
skyzh added a commit that referenced this pull request Sep 24, 2026
## Scope

Extract Week 2 Day 2 from the reviewed five-day Draft #316 onto the
merged Day 1 course. The active lesson teaches packed W4 weights and the
`quantized-matvec` checkpoint after the Day 1 capacity cache. Starter,
reference code, supplied tests, CLI, benchmarks, profiler, and
shipped-day CI expose the Day 2 contract while reserving later kernels
and checkpoint names for later PRs.

The lesson retains the signed-affine W4 derivation, eight-nibble
packing, BF16 metadata, Apple bandwidth and roofline example, matched
benchmark method, and historical Week 2 URLs. Readable text accompanies
the displayed equations. The preface identifies Day 1 as the dense
control and Day 2 as the packed path; the 4B comparison is optional.

## Local evidence

- Both learner and reference native extensions build. The Day 2
shipped-day gate collected 351 tests: 350 passed, 1 skipped. Day 1 and
Day 2 reference files passed 6/6 and 13/13. Copied learner tests collect
and stop at named TODO seams.
- Installation, documented model-free command parsing, mdBook, sitemap,
rendered local links, lint, diff and Git integrity checks passed on the
integrated local head. The two-page documentation correction was checked
again on its direct successor.
- Independent local reviews on the corrected head `fabba526` returned
factual/preservation PASS, math/command correctness GO, learner
WALKABLE, and static accessibility PASS. These verdicts are local
evidence; this public PR still requires fresh exact-head reviews.

## Open gates

Hosted macOS run 35950765030 succeeded on exact head `fabba526`: both
native extensions built and the shipped-day manifest collected 351 tests
(349 passed, 2 skipped). The runner's skip count differs from the local
author run's one skip.

This PR is Draft. Fresh public-head factual, correctness, learner, and
accessibility reviews and all current repository gates remain required
before Ready or merge. No Qwen model inference or performance benchmark
was run for this Day 2 slice, and this PR makes no speed claim. The
separate full-course Draft #316 performance screen does not validate Day
2 performance.

The current main-branch workflow deploys GitHub Pages on push. Merging
this PR would trigger publication, so the Day 2 publication decision
remains separate from the approval used for Day 1. No release, tag,
settings, visibility, or later-day work is part of this PR.

---------

Co-authored-by: Sentinel <sentinel@raft.local>
Co-authored-by: Forge <forge@raft.local>
skyzh added a commit that referenced this pull request Sep 24, 2026
## Scope

Add Week 2 Day 3 SIMD matrix prefill on top of the merged Day 2 W4
course. The active lesson introduces the `simd-matmul` checkpoint, a
10×97 partial-tile operator exercise, and a public model-output
comparison after the completed Day 1–2 checkpoints. The learner-owned
native SIMD seam remains unfinished in the starter; reference code and
supplied tests provide the runnable control. CLI, benchmark selectors,
and shipped-day CI expose Day 3 while later-day routes remain deferred.

The chapter explains the 32×32 cooperative tile, four 16×16 quadrants,
guarded edge stores, readable equation alternative, and matched
measurement method. Former Week 2 URLs remain historical. Required
measurement uses the setup-cached 0.6B model; 4B is optional. This PR
makes no speed claim.

## Local evidence

- Both learner and reference extensions build; installation checks pass.
The Day 3 shipped-day gate collected 369 tests: 368 passed, 1
optional-model skip. Direct Day 1–3 reference files passed 25 tests.
- Supplied Task 1 compares public 1×3 outputs on the fallback path. Task
2 checks the 10×97 SIMD operator against the readable packed control.
Task 3 uses fixed tiny fixtures with seeds 0 and 4, compares public 1×10
checkpoint BF16 outputs and full shapes with separate caches at
`atol=0.75`, `rtol=0.05`, and restores MLX RNG state. The unfinished
learner starter reaches its named SIMD TODO in Task 2; cumulative model
tasks still require Days 1–2 completion.
- mdBook, sitemap, rendered links, chapter command parsing, lint, diff
hygiene, and Git integrity checks passed on local head `49bda24c`. Fresh
independent local reviews returned factual/preservation PASS,
correctness/public-contract GO, learner WALKABLE, and static
accessibility PASS. These are local verdicts only.

## Open gates

Hosted macOS run 35969966892 succeeded on exact head `49bda24c`: both
native extensions built and the shipped-day manifest collected 369 tests
(367 passed, 2 skipped). The runner has one more skip than the local
author run; independent CI review will interpret the difference.

This PR is Draft. Fresh independent reviews of this exact public head,
base, and body and all current repository gates remain required before
Ready or merge. No cached Qwen benchmark was run for this Day 3 slice;
the separate full-course Draft #316 performance screen does not
establish Day 3 speed.

Chi approved normal website updates for this five-day Week 2 batch. The
ordinary main-branch Pages workflow may run after a separately gated
merge. This PR does not authorize a release, tag, settings or visibility
change, manual deployment, or later-day work.

---------

Co-authored-by: Sentinel <sentinel@raft.local>
Co-authored-by: Forge <forge@raft.local>
skyzh added a commit that referenced this pull request Sep 24, 2026
## Scope

Add Week 2 Day 4 compact model primitives on top of merged Day 3. The
active lesson, starter seams, readable reference controls, supplied
tests, CLI selectors, benchmarks, and shipped-day manifest progress
cumulatively through `rmsnorm`, `rope`, and `swiglu` after
`simd-matmul`. Learners implement the register-cached RMSNorm kernel and
its wide fallback, RoPE, and SwiGLU; Tasks 1–3 also assign the three
native CPU evaluators to report GPU-only errors. Day 5 tiled attention
and `selected` remain deferred.

The chapter preserves the RMSNorm reduction stages, RoPE head-pair math,
SwiGLU equation, cumulative model checks, and matched measurement
method. Former Week 2 URLs not reused for active Day 4 remain
historical. Required measurement uses the setup-cached Qwen3-0.6B model;
a 4B repeat is optional. This PR makes no speed claim.

## Local evidence

- The frozen tree is `a4e49d7de3c281f8ada41b5c04c0482283bcd7bc` at head
`18613b83bb4e743f1bae89bde02f8950b4947166`, directly descended from
merged Day 3 main `df51f93acf1ba24e3aa02d1460d870bd04dc7978`. The
cumulative diff is 31 Day 4 paths. The final lesson correction changed
only one page; engineering and tests retain their accepted combined
blobs.
- Both native extensions build. On the final head, independent review
passed the Day 4 reference selector (7/7), the Day 4/RMSNorm
boundary/interface set (41/41), and the unchanged extension interface
test. Copied learner tests reach named Tasks 1–3 TODOs; unfinished
cumulative model tasks remain expected learner failures. The shipped-day
manifest collects 398 nodes with Day 5 deferred.
- On the preceding combined author tree, the printed shipped-day gate
passed 397 tests with one optional-model skip. The final one-page
successor passed its focused, book, and integrity checks. Hosted macOS
run `35979136982` succeeded on this exact public head: 398 tests
collected, 396 passed, 2 skipped. The hosted log does not name the
skipped nodes.
- mdBook, sitemap, rendered local links, command syntax, diff hygiene,
and strict Git integrity passed. Four fresh independent local reviews on
the final head returned factual/preservation GO, correctness/test GO,
learner WALKABLE, and static accessibility PASS. These are local
verdicts only.

## Open gates

This PR is Draft. Hosted macOS shipped-day CI has passed on this head.
Fresh independent reviews of the exact public head, base, and body,
interpretation of the hosted skips, and current repository rules remain
gates before Ready or ordinary merge. No Day 4 performance benchmark was
run; the separate full-course Draft #316 screen does not establish a Day
4 speedup.

Chi approved normal website updates for this five-day Week 2 batch. The
ordinary main-branch Pages workflow may run after a separately gated
merge. This PR does not authorize a release, tag, settings or visibility
change, manual deployment, or later-day work.

---------

Co-authored-by: Sentinel <sentinel@raft.local>
skyzh added a commit that referenced this pull request Sep 24, 2026
## Scope

Add Week 2 Day 5 tiled dense-attention prefill and the cumulative
`selected` checkpoint on merged Day 4. Learners implement the BF16/D128
native tiled path for query length at least 9, with causal/additive
masks and grouped-query attention. Short or unsupported cases use the
readable fallback. The `selected` checkpoint exposes capacity, RMSNorm,
and tiled-attention controls independently. Required matched measurement
uses the setup-cached Qwen3-0.6B model; 4B is optional. This PR makes no
speed claim.

The lesson makes Days 1–5 active and qualifies older Week 2 URLs that
were reused by active lessons. The repaired Day 1–2 supplied tests check
public cache, quantized-operator, and packed-model behavior instead of
private layout, source text, or cache movement counters; the Day 2
lesson describes that observable scope.

## Local evidence

- Frozen head `360e485860b387ba7ecdc32bf24113cb269a7798`, tree
`7b1159092980fa82ff8a453aea259ba829747c54`, descended from merged Day 4
main `68642f05df1094ca20f92779504f737921a8dbd9`. The cumulative diff is
exactly 55 paths, each matched to an accepted source blob. The final
integration commit changes only the two repaired supplied-test files on
the corrected docs parent.
- Both native extensions and installation passed. The full unfiltered
reference suite passed 679 tests with 2 optional skips; direct Day 1–5
selectors passed 6/14/6/7/10. Focused Day 5, tiled-boundary, interface,
and repaired Day 1–2 tests passed in independent review. Copied learner
tests reach named native and preceding-day TODOs.
- mdBook/sitemap rendered 49 pages; the final-tree local scan found
1,909 links and 993 fragments with no broken targets. The 39 lesson Bash
blocks passed syntax checks, and nine focused selectors collected the
intended tests. Changed-file lint/format, diff hygiene, and strict Git
checks passed.
- Four fresh independent local reviews on this exact head returned
factual/preservation PASS, correctness/test GO, learner WALKABLE
(conditional on completing Days 1–4), and static-accessibility PASS.
These are local verdicts only. No product benchmark or real-model speed
result is claimed.

## Open gates

This PR is Draft. Hosted macOS full unfiltered reference CI succeeded on
this exact head (681 collected, 677 passed, 4 optional model skips in
Week 1 Day 5 and Week 3 Day 1); the required 0.6B tests ran. Four fresh
independent reviews of the exact public head, base, and current body
remain required before Ready or ordinary merge. The full-course Draft
#316 remains source material; its verdicts do not transfer.

Chi approved normal website updates for the five-day Week 2 batch. The
ordinary main-branch Pages workflow may run after a separately gated
merge. This PR does not authorize release, tag, settings or visibility
change, manual deployment, or later-course work.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant