@@ -200,6 +200,132 @@ unary-with-constant, and that narrowing is exactly the register in which the
200200constant, record G7 as * deliberately absent* ** in the IR's docs** , so a future
201201session does not "fix" it.
202202
203+ ** G8 — ` lzcnt_bswap_u64_to_u8 ` (tree-depth column) — NAMED 2026-09-17, not
204+ built.** The V3 facet's tail (bytes 8..16 = tiers 2–5, one aligned ` u64 ` at a
205+ compile-time offset — a PEEK, not a gather) is a ** tree path** . State the
206+ carving before the op, because "tier" and "byte" are not the same unit: the
207+ V3 facet is ` classid(4) | payload(12) ` , and ` RailSpec::v3_facet `
208+ (` src/hpc/clam_v3.rs ` ) walks ** six sequential ` (u8:u8) ` levels at key offsets
209+ ` 4..16 ` , stride 2** — so a tier is TWO bytes, one per axis, not one byte. The
210+ ` u64 ` operand at ` 8..16 ` therefore covers tiers 2–5, i.e. ** 4 tiers × 2 bytes
211+ = 8 bytes = 16 nibbles** . Each * byte* within a tier
212+ is a ** 16-ary nibble hierarchy** — OGAR canon's * "1 hex digit = 1 nibble = 1
213+ level of the 16-ary tree (` FAN_OUT=16 ` )"* — high nibble coarse. The metric on
214+ a tree path is depth-of-divergence, which at that granularity on a ` u64 ` is
215+ ` lzcnt(a ^ b) >> 2 ` — position-aware.
216+
217+ ** Which factorization of 256 this op commits to, stated because OGAR carries
218+ two and they differ by 2×.** Beside the nibble reading above, the canon also
219+ says * "256 = 4⁴ — each codebook is built as a 4-level 4-ary centroid
220+ hierarchy"* . That one is the ** Morton-interleaved** centroid space, where a
221+ nibble of ` FacetTier::morton() ` is a 2 bit × 2 bit quad-tree level across BOTH
222+ axes at once — a different tree, one level finer per step, over a different
223+ operand. G8 reads the ** raw** tail ` u64 ` , not the interleaved code, so it is
224+ the nibble tree and the shift is ` >> 2 ` (depth 0..16 over 8 bytes). A 4-ary
225+ depth over the raw bytes would be ` >> 1 ` (0..32). Both are monotone in
226+ leading-equal-bits, so ** ranking is identical either way** — the granularity
227+ only bites where depth is consumed as a * value* , which is exactly the in-cell
228+ tie-break ` x & below(depth) ` proposed below. Do not cite 4⁴ as the
229+ justification for ` >> 2 ` ; it justifies ` >> 1 ` . (Corrected 2026-09-18 after a
230+ sibling session measured the two against a fixture; the first draft of this
231+ entry asserted 4⁴ while shipping the nibble shift.) ` popcount(a ^ b) ` is ** position-blind**
232+ on it (a leaf-nibble flip counts the same as a root-nibble flip; lance-graph
233+ LATEST_STATE 2026-09-15 (8): * "popcount also finds elephant : Wal"* ), so the
234+ existing ` popcount_batch_u64 ` / ` xor_popcount ` are the wrong primitive for
235+ ranking and the right one only for the in-cell tie-break (` x & below(depth) ` ).
236+
237+ ** The byte-order wrinkle the first implementation will get backwards:** the
238+ facet is little-endian (tier 2 at the LOW byte of the tail ` u64 ` ) but the
239+ hierarchy inside a byte is MSB-coarse. ` tzcnt ` gets byte order right and
240+ nibble order wrong; ` lzcnt ` the reverse. One ` bswap ` reconciles them:
241+
242+ ``` text
243+ depth_nibbles = lzcnt(bswap(a ^ b)) >> 2 // 0..16, stepless, no branch
244+ ```
245+
246+ That is three scalar instructions (` xor ` , ` bswap ` , ` lzcnt ` ) or, as a column,
247+ ` vpshufb ` (byte reverse, ** AVX-512BW** at zmm width) + ` vplzcntq `
248+ (** AVX-512CD** ) over 8 rows per zmm — BOTH feature bits, not CD alone; a
249+ target with CD but not BW needs a different byte reversal. Per 8-row group
250+ there is ** no operand padding and no tail descent** — eight tails tile a zmm
251+ exactly, which is
252+ the whole reason the ` u64 ` width is the right one here (and consistent with
253+ blackboard (18): the tail optimisation is inert on power-of-two populations
254+ because there is no tail; stripping HEEL/HIP makes the operand * start*
255+ power-of-two). ** That is a claim about the OPERAND, not the population:** a
256+ row count not divisible by 8 still leaves a column remainder, exactly as the
257+ slice-level ` U64x8 ` ops handle with ` pad_tail ` . No multiple-of-eight input
258+ contract is stated or intended, so an implementation keeps a remainder path
259+ and tests non-multiple-of-eight lengths. What is eliminated is the
260+ * within-operand* padding a narrower width would need, not the last partial
261+ group. Consumer: basin-local similarity in lance-graph (` FacetCascade `
262+ tail) and the CAKES nearest-ranking on ` NiblePath ` (lance-graph
263+ ` ISS-NIBLEPATH-FOLD-IS-CARRIER-2-UNMASKED ` , the packed-` u64 ` carrier where the
264+ same ` lzcnt ` IS the fold). Realizations: ` vplzcntq ` on avx512cd; avx2 has no
265+ vector lzcnt — polyfill via the float-exponent trick or a scalar peel (measure
266+ which); NEON has ` clz ` on 32-bit lanes (two halves + select); wasm/scalar flat;
267+ and the ** nightly** ` core::simd ` arm (` src/simd_nightly/ ` ) — six in total,
268+ which is the count the falsifier below means.
269+ Pre-registered falsifier: the column op must equal the scalar
270+ ` (a ^ b).swap_bytes().leading_zeros() >> 2 ` on every row of a 64k fixture at
271+ all six realizations, AND a disable-run without the ` bswap ` must fail on a
272+ fixture whose divergence is in a low nibble of a high tier (that is the
273+ backwards-implementation trap, and the test must be able to see it).
274+ ** Plus a third arm, because the first two cannot catch a wrong shift:** both
275+ compare the column op against ` (a ^ b).swap_bytes().leading_zeros() >> 2 ` ,
276+ which is the same formula in a different spelling, so they agree even if the
277+ shift is wrong. The third arm builds the fixture * independently* — plant a
278+ divergence at a KNOWN tier and nibble, assert the KNOWN depth — and is the
279+ only one that discriminates ` >> 2 ` from ` >> 1 ` . ** Measured** (2026-09-18, all
280+ 16 positions, against a scalar model): for tier ` t ` ∈ 2..5, byte-in-tier ` b `
281+ (0 = ` lo ` , the LOWER address; 1 = ` hi ` ) and nibble-in-byte ` k ` (0 = high/MSB,
282+ 1 = low/LSB),
283+
284+ ```
285+ depth = 4*(t-2) + 2*b + k // 0..15; 16 iff the tails are identical
286+ ```
287+
288+ — exact at every one of the 16 positions. (A first draft of this arm wrote
289+ ` 2*(t-2) + n ` , which is right only for tier 2 and drifts by ` 2*(t-2) `
290+ thereafter; each tier is 2 bytes = ** 4** nibbles, not 2. It was caught by
291+ running the model, which is the point of the arm.)
292+
293+ ** Caveat the measurement surfaced, and the reason the consumer gate matters
294+ more than it looked.** ` FacetTier ` is ` { lo, hi } ` with ` lo ` at the LOWER
295+ address, so after the ` bswap ` a single ` lzcnt ` over the raw tail walks
296+ ` lo ` BEFORE ` hi ` within every tier — i.e. it alternates between the two
297+ chains every two nibbles. But ` hi_chain ` and ` lo_chain ` are documented as two
298+ ** orthogonal** hierarchies (` facet.rs ` : * "the ` hi ` chain prefix-routes one
299+ hierarchy, the ` lo ` chain the orthogonal one"* ), not coarse and fine of one.
300+ So raw-tail depth is a ** mixed-axis** metric: still a valid monotone
301+ tie-breaker, NOT a depth in either hierarchy. A consumer that wants one
302+ axis needs a per-chain fold (the existing ` hi_distance ` / ` lo_distance ` ) or a
303+ deinterleave before the ` lzcnt ` .
304+
305+ ** And ` clam_v3.rs ` already answers which carving the real bake wants — it is
306+ not the pair reading.** ` RailSpec ` carries TWO carvings, and the doc comment
307+ records a measurement against them: the interleaved ` X:Y ` pair reading
308+ (` v3_facet ` , 6 levels, stride 2) * "fits only 44.25 % of paths"* on the
309+ medcare bake, while the contiguous per-axis slab (` RailSpec::slab ` , 12 levels,
310+ stride 1) * "fits 99.62 % in twelve levels"* . That inverts the gate for G8: on
311+ the ** slab** carving a contiguous ` u64 ` IS one axis, the mixed-axis caveat
312+ above evaporates, and a plain ` lzcnt(bswap(·)) ` is exactly right. On the
313+ ** pair** carving it is a mixed-axis tiebreaker. So G8's consumer gate must
314+ name the CARVING as well as the call site, and the measured 99.62 %/44.25 %
315+ split says the slab is the likelier target. Do not build against the pair
316+ reading on the strength of it being the zero-fallback default.
317+
318+ ** Stacking, not widening, is how the levels grow.** The same module: * "a class
319+ that needs more than six levels does not widen a byte — it stacks a second
320+ register, e.g. into the edge lane, and chains it (` RailSpec::stacked ` ). Depth
321+ then runs 0..=12 over two registers, same hole rule, same arithmetic."* The
322+ edge lane at ` 16..32 ` is explicitly contemplated as that continuation register
323+ (* "` 16..32 ` may be a continuation register (if the spec says so)"* ). A G8
324+ column op over a stacked pair is therefore two operands chained, never one
325+ wider one — which is also why no ` u128 ` variant of G8 is proposed. Gate
326+ before building: one named consumer call site that ranks by depth — the same
327+ count rule G5 carries.
328+
203329** The nightly arm is AHEAD of the stable arms, and it is the contract
204330reference.** ` src/simd_nightly/ ` carries ** 18 compare-to-mask pairs across
205331every width** ; the stable arms have a subset. So N2/N3 were not adding a
0 commit comments