Skip to content

Commit 1bc8673

Browse files
authored
Merge pull request #311 from AdaWorldAPI/claude/c64-6502-falsifier-shztkk
ternlogq tail: padded zmm vs zmm→ymm→xmm descent, measured (AVX-512)
2 parents c746735 + f164f53 commit 1bc8673

6 files changed

Lines changed: 1071 additions & 1 deletion

File tree

.claude/blackboard.md

Lines changed: 108 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,111 @@
1+
## 2026-09-16 (13) — the `VPTERNLOGQ` tail is a DESCENT, not a pad (5–8×); a 64×2 re-apply on a full-width mask is NOT (0.5–0.7×); the GEMM block-stop tail is INERT (0.99–1.02×)
2+
3+
Three probes, one question in three places (operator: *"instead of padding the
4+
tail you could simply split the N×64×8 + N×64×2 tail"*, then *"would the gather
5+
vs re-apply be faster with 64×2 instead of 64×8"*, then *"would the MKL GEMM
6+
stop logic profit from a tail optimization"*). All AVX-512 (`.cargo/config-v4.toml`,
7+
`avx512f=true` printed by each program), release, `black_box` on inputs AND
8+
outputs, every arm gated bit-identical before timing. Branch
9+
`claude/c64-6502-falsifier-shztkk`, PR #311.
10+
11+
### 1. `examples/ternlogq_tail_descent_probe.rs` — the TAIL. YES.
12+
13+
`mask_ternlog` chunks over `U64x8` (512 rows) and ends on three `pad_tail`s
14+
(zero-fill three `[u64; 8]`, one zmm op, prefix copy). The alternative: descend
15+
zmm → ymm → xmm (`4 + 2 + 1`, every lane live, all in vector registers).
16+
17+
Crux — for a remainder of `t` words, tail only, ns/call, 3 runs:
18+
19+
| t | shape | P padded zmm | **G greedy** | X all-xmm | winner |
20+
|---:|---|---:|---:|---:|---|
21+
| 1 | `1` | 13.1–13.8 | 1.78–1.90 | 1.49–1.60 | G = X (identical code) |
22+
| 2 | `2` | 13.0–13.7 | **1.74–1.94** | 2.57–2.78 | G |
23+
| 3 | `2+1` | 17.1–18.6 | **2.13–2.36** | 2.82–3.23 | G |
24+
| 4 | `4` vs `2+2` | 12.8–13.6 | **1.59–1.76** | 2.92–3.00 | G |
25+
| 5 | `4+1` vs `2+2+1` | 17.3–18.3 | **2.08–2.13** | 3.23–3.49 | G |
26+
| **6** | **`4+2`** vs `2+2+2` | 17.0–18.3 | **2.26–2.74** | 3.29–3.68 | **G** |
27+
| 7 | `4+2+1` vs `2+2+2+1` | 16.8–18.5 | **2.38–2.57** | 3.68–3.95 | G |
28+
29+
Greedy widest-first wins every `t ≥ 2`: 1.3–1.6× over all-xmm, 5–8× over
30+
padding. The operator's crux (`t = 6`): one ymm + one xmm beats three xmm.
31+
32+
End to end through the real chunk loop:
33+
34+
| words | rows | tail | padded ns | descend ns | ratio |
35+
|---:|---:|---:|---:|---:|---:|
36+
| 3 (`ogar-r2il` `CallMask`) | 192 | 3 | 18.1 | 2.3 | **7.7–8.0×** |
37+
| 1–7 | 64–448 | 1–7 | 13–23 | 2.0–2.6 | 5.9–8.0× |
38+
| 9 | 576 | 1 | 14.2 | 3.7 | 4.2× |
39+
| 11 | 704 | 3 | 21.9 | 3.6 | 5.4–5.6× |
40+
| 31 | 1 984 | 7 | 23.6 | 6.8 | 2.5–3.3× |
41+
| 194 | 12 416 | 2 | 41.4 | 26.5 | 1.3–1.4× |
42+
| 8 / 16 / 24 / 64 || none ||| 1.0–1.4× (loop shape, not tail) |
43+
44+
asm: 33 zmm + 3 ymm + 6 xmm `vpternlogq`, folded memory operands, **zero** GPR
45+
and/or/xor on lane data — a descent is not the scalar peel
46+
`scripts/codegen-witness.sh` caps at `SLICE_GPR_CAP=6`. Follow-up named, not
47+
built: `U64x4::ternlog` / `U64x2::ternlog` on the facade + rewire
48+
`mask_ternlog`'s tail; an un-gated `pack<const L>` sibling of `pack_under` to
49+
retire the 12 hand-rolled `if !tail.is_empty()` sites.
50+
51+
### 2. `examples/ternlogq_sparse_reapply_probe.rs` — FULL-WIDTH sparse frontier. NO.
52+
53+
1 024 words (65 536 rows, the MQ / `lgj_hop` population), `dst = src ∧ gate ∧
54+
elig`, arms: F8 full zmm pass (shipped), S8/S4/S2 chunk-skip at zmm/ymm/xmm
55+
(`vptestmq` → kortest), S1 GPR floor, W per-bit gather walk. ns/call, 3 runs:
56+
57+
| frontier | shape | live | F8 | S8 | S4 | S2 | S1 | W |
58+
|---|---|---:|---:|---:|---:|---:|---:|---:|
59+
| 0.01 % | uniform | 7 | 143–164 | **121–141** | 148–168 | 239–277 | 192–233 | 404–435 |
60+
| 0.01 % | clustered | 7 | 148–163 | **125–136** | 144–145 | 240–242 | 223 | 358–368 |
61+
| 0.1 % | clustered | 66 | 127–161 | **103–132** | 137–145 | 227–240 | 198–223 | 394–421 |
62+
| 1 % | uniform | 655 | **153–162** | 190–197 | 216–222 | 277–289 | 198–224 | 918–1 013 |
63+
| 1 % | clustered | 655 | 152–162 | **123–132** | 129–146 | 190–240 | 223–227 | 933–981 |
64+
| 10 % | uniform | 6 554 | **153–162** | 188–200 | 213–222 | 314–334 | 216–224 | 9 987–10 849 |
65+
| 10 % | clustered | 6 554 | 152–161 | **135–141** | 144–157 | 239–258 | 193–222 | 6 187–6 926 |
66+
| 100 % | either | 65 536 | **145–161** | 172–205 | 194–222 | 313–339 | 201–229 | 52–58 k |
67+
68+
S2 is 1.5–2.1× SLOWER than the full pass everywhere; S4 never beats S8; the
69+
skip is worth ≤ 1.24× on clustered frontiers and LOSES on uniform ≥ 1 %
70+
(branch mispredicts). 150 ns for 32 KiB of traffic is L1 bandwidth; narrower
71+
chunks are more iterations, not less work. The gather walks all 1 024 words
72+
before knowing they are empty (≥ 360 ns at 7 bits, ~0.8 ns/bit after). The
73+
64×2 rung is a tail instrument only.
74+
75+
### 3. `kernels_avx512.rs::block_stop_probe` (ignored test) — the GEMM stop. INERT.
76+
77+
`sgemm_blocked`'s M-stop pads the last `MR=6` panel and the ukernel computes all
78+
six accumulators regardless of `mr_eff`; every power-of-two `m` has such a tail
79+
(128 = 21·6+2, 256 = 42·6+4, 512 = 85·6+2, 1024 = 170·6+4). A test-local `R×16`
80+
tail ukernel (R ∈ {2, 4}) on the tail tile only, bit-identical, best of 9:
81+
82+
| m×n×k | tail | shipped ms | desc ms | ratio | FMA waste |
83+
|---|---:|---:|---:|---:|---:|
84+
| 126×256×256 | 0 | 0.183 | 0.185 | 0.987× | 0.0 % |
85+
| 128×256×256 | 2 | 0.190 | 0.192 | 0.989× | 3.1 % |
86+
| 130×256×256 | 4 | 0.190 | 0.191 | 0.992× | 1.5 % |
87+
| 132×256×256 | 0 | 0.193 | 0.190 | 1.012× | 0.0 % |
88+
| 128³ | 2 | 0.056 | 0.055 | 1.015× | 3.1 % |
89+
| 256³ | 4 | 0.347 | 0.350 | 0.992× | 0.8 % |
90+
| 512³ | 2 | 3.110 | 3.054 | 1.018× | 0.8 % |
91+
| 1024³ | 4 | 32.57 | 32.04 | 1.016× | 0.2 % |
92+
93+
Noise. ~66 GFLOP/s at 1024³ (half of one core's FMA peak): packing and memory
94+
traffic hide the padded rows. K-stop has no waste; N-stop padding is on lanes
95+
the FMA unit processes anyway — no lane-width descent applies to GEMM. The BF16
96+
`vdpbf16ps` path's stop problem is ONE accumulator chain per row
97+
(`amx_matmul.rs:686-697`, latency-bound), not its tails.
98+
99+
### The rule the three share
100+
101+
The tail descent paid where the pad cost **loads and a copy** (3 zero-fills +
102+
prefix copy per call). It buys nothing where the padding is **ALU on lanes the
103+
unit processes anyway** (GEMM N-stop, sparse re-apply, GEMM M-stop hidden
104+
behind bandwidth). Sibling finding the same day, lance-graph #1241: the facet's
105+
per-axis LCP was gathering + re-folding per call; reading the single `u128`
106+
register masked to the axis bytes (the `-f` done ONCE at mint) took it
107+
12.5 → 5.8 ns for both axes.
108+
1109
## 2026-09-16 (12) — ⊘ the G1 ratio was INFLATED by dead-store elimination; corrected 6.75× → 5.84×, and the conclusion survives
2110

3111
codex P2 on PR #309, and it was load-bearing. Entry (11)'s numbers are

Cargo.toml

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -63,6 +63,14 @@ required-features = ["std"]
6363
name = "r2il_column_scan_probe"
6464
required-features = ["std"]
6565

66+
[[example]]
67+
name = "ternlogq_tail_descent_probe"
68+
required-features = ["std"]
69+
70+
[[example]]
71+
name = "ternlogq_sparse_reapply_probe"
72+
required-features = ["std"]
73+
6674
[[example]]
6775
name = "hex_tenant_mq_probe"
6876
required-features = ["std"]

0 commit comments

Comments
 (0)