Skip to content

Commit e1ef350

Browse files
authored
Merge pull request #317 from AdaWorldAPI/claude/great-pascal-k96kok
masking-ops: name G8 — the tree-depth column `lzcnt(bswap(x)) >> 2`; popcount is its tie-break
2 parents 6f40bc6 + 40a71ad commit e1ef350

3 files changed

Lines changed: 177 additions & 0 deletions

File tree

‎.claude/blackboard.md‎

Lines changed: 18 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,21 @@
1+
## 2026-09-17 (19) — G8 named: a tree-depth column (`lzcnt(bswap(x)) >> 2`) is the missing primitive for basin-local ranking; popcount is only its tie-break
2+
3+
Filed, not built. Full text in `masking-ops-state.md` § OUTLOOK G8 and the
4+
2026-09-17 EPIPHANIES entry. Two things for the next session:
5+
6+
- **Loose end:** the gate is a named consumer call site that ranks by depth.
7+
Candidates: lance-graph `FacetCascade` tail ranking (basin-local similarity)
8+
and `NiblePath::common_prefix_depth` (the packed-`u64` carrier where this
9+
`lzcnt` IS the fold — `ISS-NIBLEPATH-FOLD-IS-CARRIER-2-UNMASKED`). Either
10+
one, once it exists as a call, licenses the build.
11+
- **Decision recorded:** do NOT add a `popcount` variant for this; the existing
12+
`popcount_batch_u64` already serves the tie-break. The gap is `lzcnt` (+
13+
`bswap`), and (8) on the lance-graph board had already named the `lzcnt`
14+
half.
15+
16+
Consistent with (18): stripping HEEL/HIP makes the operand start power-of-two,
17+
so the tail-padding question never arises for this op.
18+
119
## 2026-09-16 (18) — the tail OPTIMISATION IS INERT ON EVERY POWER-OF-TWO POPULATION >= 512 ROWS, including the 4096-row tile
220

321
Established before writing the rewrite (15)-(17) argued for, and it changes

‎.claude/board/EPIPHANIES.md‎

Lines changed: 33 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,38 @@
11
# ndarray — Epiphanies (append-only)
22

3+
## 2026-09-17 — On a tree-path register the metric is lzcnt, popcount is the tie-break, and little-endian bytes need one bswap first
4+
**Status:** FINDING (the layout and the instruction facts) + NAMED GAP (G8, not built)
5+
**Scope:** @simd-savant @family-codec-smith domain:masking-ops domain:v3-facet
6+
**Cross-ref:** `.claude/knowledge/masking-ops-state.md` § OUTLOOK G8; lance-graph
7+
`E-THREE-CARRIERS-THREE-FOLDS-1`, `ISS-NIBLEPATH-FOLD-IS-CARRIER-2-UNMASKED`,
8+
LATEST_STATE 2026-09-15 (6)–(8); blackboard (18)
9+
10+
Within a basin every V3 facet shares HEEL and HIP by construction, so the
11+
6-tier LCP saturates and the information is in the tail — bytes 8..16, tiers
12+
2–5, one aligned `u64` at a compile-time offset. Loading it is a PEEK (an
13+
address), not a mask (a reassembly), so register ops on it are legitimate
14+
under the carrier doctrine; and stripping the shared prefix is a load-offset
15+
choice that costs nothing. Two consequences the workspace had not written down:
16+
17+
1. **Padding is useless once you can strip.** The tail descent / `[u64; 8]`
18+
zero-padding story exists because a 128-bit facet operand is ragged in a
19+
512-bit register. A 64-bit tail is not: eight tails tile a zmm exactly.
20+
Blackboard (18) found the tail optimisation inert on power-of-two
21+
populations *because there is no tail*; this is the same fact from the
22+
operand side — choose the width that starts power-of-two.
23+
2. **The op on it is `lzcnt`, not `popcount`.** Each tier byte is a 4-level
24+
4-ary centroid tree, so the tail is a path and its metric is depth of
25+
divergence. Popcount is position-blind on a path (leaf flip == root flip).
26+
Ranking = `lzcnt(bswap(a ^ b)) >> 2`; popcount belongs only to the
27+
tie-break among equal-depth candidates (`x & below(depth)`) — coarse by
28+
depth, fine by density.
29+
30+
The `bswap` is load-bearing and is the trap: LE byte order puts the root tier
31+
at the LOW byte, but nibble order inside a byte is MSB-coarse, so neither
32+
`tzcnt` nor `lzcnt` alone reads the path in one direction. G8 records the
33+
falsifier that catches a `bswap`-less implementation. Gate before building is
34+
the G5 rule: one named consumer call site.
35+
336
## 2026-07-29 — Which PACKAGE pulls a dep decides whether it can consume you
437
**Status:** FINDING
538
**Scope:** @simd-savant @truth-architect domain:build-graph

‎.claude/knowledge/masking-ops-state.md‎

Lines changed: 126 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -200,6 +200,132 @@ unary-with-constant, and that narrowing is exactly the register in which the
200200
constant, record G7 as *deliberately absent* **in the IR's docs**, so a future
201201
session does not "fix" it.
202202

203+
**G8 — `lzcnt_bswap_u64_to_u8` (tree-depth column) — NAMED 2026-09-17, not
204+
built.** The V3 facet's tail (bytes 8..16 = tiers 2–5, one aligned `u64` at a
205+
compile-time offset — a PEEK, not a gather) is a **tree path**. State the
206+
carving before the op, because "tier" and "byte" are not the same unit: the
207+
V3 facet is `classid(4) | payload(12)`, and `RailSpec::v3_facet`
208+
(`src/hpc/clam_v3.rs`) walks **six sequential `(u8:u8)` levels at key offsets
209+
`4..16`, stride 2** — so a tier is TWO bytes, one per axis, not one byte. The
210+
`u64` operand at `8..16` therefore covers tiers 2–5, i.e. **4 tiers × 2 bytes
211+
= 8 bytes = 16 nibbles**. Each *byte* within a tier
212+
is a **16-ary nibble hierarchy** — OGAR canon's *"1 hex digit = 1 nibble = 1
213+
level of the 16-ary tree (`FAN_OUT=16`)"* — high nibble coarse. The metric on
214+
a tree path is depth-of-divergence, which at that granularity on a `u64` is
215+
`lzcnt(a ^ b) >> 2` — position-aware.
216+
217+
**Which factorization of 256 this op commits to, stated because OGAR carries
218+
two and they differ by 2×.** Beside the nibble reading above, the canon also
219+
says *"256 = 4⁴ — each codebook is built as a 4-level 4-ary centroid
220+
hierarchy"*. That one is the **Morton-interleaved** centroid space, where a
221+
nibble of `FacetTier::morton()` is a 2 bit × 2 bit quad-tree level across BOTH
222+
axes at once — a different tree, one level finer per step, over a different
223+
operand. G8 reads the **raw** tail `u64`, not the interleaved code, so it is
224+
the nibble tree and the shift is `>> 2` (depth 0..16 over 8 bytes). A 4-ary
225+
depth over the raw bytes would be `>> 1` (0..32). Both are monotone in
226+
leading-equal-bits, so **ranking is identical either way** — the granularity
227+
only bites where depth is consumed as a *value*, which is exactly the in-cell
228+
tie-break `x & below(depth)` proposed below. Do not cite 4⁴ as the
229+
justification for `>> 2`; it justifies `>> 1`. (Corrected 2026-09-18 after a
230+
sibling session measured the two against a fixture; the first draft of this
231+
entry asserted 4⁴ while shipping the nibble shift.) `popcount(a ^ b)` is **position-blind**
232+
on it (a leaf-nibble flip counts the same as a root-nibble flip; lance-graph
233+
LATEST_STATE 2026-09-15 (8): *"popcount also finds elephant : Wal"*), so the
234+
existing `popcount_batch_u64` / `xor_popcount` are the wrong primitive for
235+
ranking and the right one only for the in-cell tie-break (`x & below(depth)`).
236+
237+
**The byte-order wrinkle the first implementation will get backwards:** the
238+
facet is little-endian (tier 2 at the LOW byte of the tail `u64`) but the
239+
hierarchy inside a byte is MSB-coarse. `tzcnt` gets byte order right and
240+
nibble order wrong; `lzcnt` the reverse. One `bswap` reconciles them:
241+
242+
```text
243+
depth_nibbles = lzcnt(bswap(a ^ b)) >> 2 // 0..16, stepless, no branch
244+
```
245+
246+
That is three scalar instructions (`xor`, `bswap`, `lzcnt`) or, as a column,
247+
`vpshufb` (byte reverse, **AVX-512BW** at zmm width) + `vplzcntq`
248+
(**AVX-512CD**) over 8 rows per zmm — BOTH feature bits, not CD alone; a
249+
target with CD but not BW needs a different byte reversal. Per 8-row group
250+
there is **no operand padding and no tail descent** — eight tails tile a zmm
251+
exactly, which is
252+
the whole reason the `u64` width is the right one here (and consistent with
253+
blackboard (18): the tail optimisation is inert on power-of-two populations
254+
because there is no tail; stripping HEEL/HIP makes the operand *start*
255+
power-of-two). **That is a claim about the OPERAND, not the population:** a
256+
row count not divisible by 8 still leaves a column remainder, exactly as the
257+
slice-level `U64x8` ops handle with `pad_tail`. No multiple-of-eight input
258+
contract is stated or intended, so an implementation keeps a remainder path
259+
and tests non-multiple-of-eight lengths. What is eliminated is the
260+
*within-operand* padding a narrower width would need, not the last partial
261+
group. Consumer: basin-local similarity in lance-graph (`FacetCascade`
262+
tail) and the CAKES nearest-ranking on `NiblePath` (lance-graph
263+
`ISS-NIBLEPATH-FOLD-IS-CARRIER-2-UNMASKED`, the packed-`u64` carrier where the
264+
same `lzcnt` IS the fold). Realizations: `vplzcntq` on avx512cd; avx2 has no
265+
vector lzcnt — polyfill via the float-exponent trick or a scalar peel (measure
266+
which); NEON has `clz` on 32-bit lanes (two halves + select); wasm/scalar flat;
267+
and the **nightly** `core::simd` arm (`src/simd_nightly/`) — six in total,
268+
which is the count the falsifier below means.
269+
Pre-registered falsifier: the column op must equal the scalar
270+
`(a ^ b).swap_bytes().leading_zeros() >> 2` on every row of a 64k fixture at
271+
all six realizations, AND a disable-run without the `bswap` must fail on a
272+
fixture whose divergence is in a low nibble of a high tier (that is the
273+
backwards-implementation trap, and the test must be able to see it).
274+
**Plus a third arm, because the first two cannot catch a wrong shift:** both
275+
compare the column op against `(a ^ b).swap_bytes().leading_zeros() >> 2`,
276+
which is the same formula in a different spelling, so they agree even if the
277+
shift is wrong. The third arm builds the fixture *independently* — plant a
278+
divergence at a KNOWN tier and nibble, assert the KNOWN depth — and is the
279+
only one that discriminates `>> 2` from `>> 1`. **Measured** (2026-09-18, all
280+
16 positions, against a scalar model): for tier `t` ∈ 2..5, byte-in-tier `b`
281+
(0 = `lo`, the LOWER address; 1 = `hi`) and nibble-in-byte `k` (0 = high/MSB,
282+
1 = low/LSB),
283+
284+
```
285+
depth = 4*(t-2) + 2*b + k // 0..15; 16 iff the tails are identical
286+
```
287+
288+
— exact at every one of the 16 positions. (A first draft of this arm wrote
289+
`2*(t-2) + n`, which is right only for tier 2 and drifts by `2*(t-2)`
290+
thereafter; each tier is 2 bytes = **4** nibbles, not 2. It was caught by
291+
running the model, which is the point of the arm.)
292+
293+
**Caveat the measurement surfaced, and the reason the consumer gate matters
294+
more than it looked.** `FacetTier` is `{ lo, hi }` with `lo` at the LOWER
295+
address, so after the `bswap` a single `lzcnt` over the raw tail walks
296+
`lo` BEFORE `hi` within every tier — i.e. it alternates between the two
297+
chains every two nibbles. But `hi_chain` and `lo_chain` are documented as two
298+
**orthogonal** hierarchies (`facet.rs`: *"the `hi` chain prefix-routes one
299+
hierarchy, the `lo` chain the orthogonal one"*), not coarse and fine of one.
300+
So raw-tail depth is a **mixed-axis** metric: still a valid monotone
301+
tie-breaker, NOT a depth in either hierarchy. A consumer that wants one
302+
axis needs a per-chain fold (the existing `hi_distance` / `lo_distance`) or a
303+
deinterleave before the `lzcnt`.
304+
305+
**And `clam_v3.rs` already answers which carving the real bake wants — it is
306+
not the pair reading.** `RailSpec` carries TWO carvings, and the doc comment
307+
records a measurement against them: the interleaved `X:Y` pair reading
308+
(`v3_facet`, 6 levels, stride 2) *"fits only 44.25 % of paths"* on the
309+
medcare bake, while the contiguous per-axis slab (`RailSpec::slab`, 12 levels,
310+
stride 1) *"fits 99.62 % in twelve levels"*. That inverts the gate for G8: on
311+
the **slab** carving a contiguous `u64` IS one axis, the mixed-axis caveat
312+
above evaporates, and a plain `lzcnt(bswap(·))` is exactly right. On the
313+
**pair** carving it is a mixed-axis tiebreaker. So G8's consumer gate must
314+
name the CARVING as well as the call site, and the measured 99.62 %/44.25 %
315+
split says the slab is the likelier target. Do not build against the pair
316+
reading on the strength of it being the zero-fallback default.
317+
318+
**Stacking, not widening, is how the levels grow.** The same module: *"a class
319+
that needs more than six levels does not widen a byte — it stacks a second
320+
register, e.g. into the edge lane, and chains it (`RailSpec::stacked`). Depth
321+
then runs 0..=12 over two registers, same hole rule, same arithmetic."* The
322+
edge lane at `16..32` is explicitly contemplated as that continuation register
323+
(*"`16..32` may be a continuation register (if the spec says so)"*). A G8
324+
column op over a stacked pair is therefore two operands chained, never one
325+
wider one — which is also why no `u128` variant of G8 is proposed. Gate
326+
before building: one named consumer call site that ranks by depth — the same
327+
count rule G5 carries.
328+
203329
**The nightly arm is AHEAD of the stable arms, and it is the contract
204330
reference.** `src/simd_nightly/` carries **18 compare-to-mask pairs across
205331
every width**; the stable arms have a subset. So N2/N3 were not adding a

0 commit comments

Comments
 (0)