Skip to content

zspace: Fisher-Z entry tax without materializing z (ZGamma codes) + binomial Hamming null; batch Welford recovered from #323 - #325

Merged
AdaWorldAPI merged 8 commits into
masterfrom
claude/llvm-codegen-polyfill-gni3cw
Sep 24, 2026
Merged

AdaWorldAPI merged 8 commits into
masterfrom
claude/llvm-codegen-polyfill-gni3cw

Conversation

@AdaWorldAPI

@AdaWorldAPI AdaWorldAPI commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Why

Operator rules:

  • Anything cosine-shaped pays the Fisher-Z entry tax before it is thresholded, averaged, banded or ranked by margin. There are no exceptions, including lab and calibration code.
  • Anything bitpacked goes through HDR popcount stacking on the Hamming distance itself.
  • Fisher-Z is never materialized. A value is normalized once and stays in that space.
  • Bit-exact on every target.

Before this PR, ndarray had no Fisher transform at all.

This PR also carries two commits that were meant for #323 but missed it: #323 was merged at edbd97f, one push before them. They are cherry-picked here unchanged: moments_u32 + Cascade::observe_batch, and their re-export.

What

hpc::zspace (7bb45fb → bb14de6 → fd45c89 → f2a461a)

ZGamma { z_min, z_range } is the z-space envelope: 8 little-endian bytes.

  • fit(cosines) scans only the cosine min and max. Fisher-Z is monotone, so no z value is ever stored.
  • encode / encode_batch go straight from cosine to an i8 code, and only the code is written. Both call one shared quantize().
  • The code: ((z−z_min)/z_range)·254 − 127, rounded ties-to-even and saturated to the symmetric range −127..=127. −128 is reserved as ZGamma::NAN_CODE.
  • threshold_code(c) converts a threshold into code space once. There is no decode back to cosine.

Bit-exact across targets (f2a461a). simd_ln_f32 is a per-lane libm .ln(), so the first version inherited the platform libm. Measured over an 8,539-value grid:

f32 z f64 z
values differing, x86-64 glibc vs wasm32 624 566
size of the difference 1–2 ulp 1–2 ulp

The codes happened to agree on that grid, but a one-ulp difference at a rounding boundary flips a code.

ln_det replaces libm in the code path. It is musl/FreeBSD logf, built only from IEEE add, sub, mul, div and bit operations in a fixed order. Rust never contracts these into FMA, so every conforming target gets the same bits. Its error is under 1 ulp over its domain [2⁻²⁴, 2). Its x86 z digest now equals the digest wasm's (musl-derived) libm produced before.

Where this departs from bgz-tensor. It shares FamilyGamma's byte layout, but not its code. FamilyGamma predates this analysis and was not calibrated for exactness:

FamilyGamma behaviour measured effect
truncation toward zero the code-0 bucket is twice as wide as the others, and every code is pulled half a step inward (median bias: +0.49 / −0.49). Rounding: +0.006 / −0.002.
−128 used as a data level asymmetric range
NaN mapped to 0 collides with a real code
platform libm ln not bit-exact across targets

Unchanged:

  • fisher_z / hyperbolic_depth stay, for single f64 scalars only (a report statistic, a threshold). They keep libm and are documented as not bit-exact.
  • hamming_null_z stays. It is the binomial-null z for bitpacked data.

Removed:

  • fisher_z_f32_batch wrote a buffer of z values.
  • fisher_z_inv was the way back to cosine.

FidelityReport pays the tax (b8babbd)

  • Adds pearson_z / spearman_z / cosine_z (scalar statistics).

Batch Welford, recovered (90f1128, 414aa8e)

  • moments_u32 computes exact integer (n, Σx, Σx²) through U64x8.
  • MomentsU32::merge is exact and order-independent.
  • Cascade::observe_batch uses the Chan–Golub–LeVeque merge.

wasm enforcement (aa3dfd7)

CI never runs cargo test on wasm. So crates/wasm-simd-parity's node selfcheck now rebuilds the grid and checks the pinned codes digest (rc 0x501) and batch == scalar (rc 0x502).

Tests

zspace (14 tests):

  • the 12 earlier tests: bit identity, symmetric saturation, spacing in z, degenerate inputs, threshold, round-trip through 8 LE bytes, no median bias, NaN sentinel, the Fisher-Z and binomial pins;
  • ln_det is within 1 ulp across its domain, sampling every 257th bit pattern;
  • pinned digests (GOLDEN_CODES, GOLDEN_Z32) that must hold on every target.

Disable checks — each deliberate bug makes the listed test fail:

deliberate bug what fails
truncate instead of round median-bias test
saturate to −128 3 tests
NaN → 0 sentinel test
clamp removed from the batch lanes bit-identity test
Fisher transform removed 3 tests
degenerate-range guard removed degenerate-population test
libm .ln() back in the scalar path pinned z digest
altered wasm golden constant selfcheck returns rc 0x501

Test runs by build:

build result
config-v3 (AVX2) 14/14
config-v4 (AVX-512) 14/14
aarch64 NEON under qemu 14/14
wasm32 simd128 under node selfcheck OK (codes digest)

The doc tests pass. clippy -D warnings and fmt are clean.

Open questions (not decided in this PR)

  • What does code 0 mean? Here code 0 is the midpoint of the fitted z range, not zero correlation. The alternative is a scale-only envelope, code = round(127·z/z_absmax): the sign of a code becomes the sign of the evidence. It costs half the resolution when a population is one-sided.
  • HDR popcount stacking vs Fisher-Z. Bitpacked data stays in Hamming, measured against the binomial null, and is never converted Hamming → cosine → Fisher-Z. The two scores are comparable only in σ units: fisher_z·√(n−3) and hamming_null_z are both ~N(0,1) under their nulls.
  • simd_ln_f32 itself is still per-lane libm, used by callers outside zspace. Whether the facade's ln should become ln_det (bit-exact, and vectorizable in the same op order) is a separate decision.

Not in this PR

  • lance-graph sites that skip the Fisher-Z tax:
    • the raw SQL cosine UDFs;
    • the contract's similarity_z / fisher_z_inverse.
  • Rim-clamp disagreements between Fisher-Z copies.
  • helix's hardcoded γ.
  • bgz-tensor FamilyGamma: truncation, −128, NaN and libm are all present there too. Fixing it changes stored codes, so it needs its own decision.

🤖 Generated with Claude Code

https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm

Summary by CodeRabbit

  • New Features
    • Added Fisher-Z score views for Pearson, Spearman, and cosine similarity results, making these values available for consistent comparison and thresholding.
    • Added statistical scoring for bitpacked Hamming distances and compact encoding of cosine similarity values, including batch encoding and threshold lookup.
    • Added batch observation support that combines distance measurements and can flag significant shifts in the observed mean.
    • Added utilities for calculating and combining exact moments—count, mean, and variance—from unsigned integer measurements.

`moments_u32(&[u32]) -> MomentsU32 { n, sum, sum_sq }`: exact integer
moments, eight lanes at a time through U64x8. Squares are split into
32-bit halves so every lane add stays below 2^32; lanes drain into u128
totals every 2^28 chunks, before an 8-lane reduce_sum could overflow u64.
`MomentsU32::merge` is plain integer addition — exact, associative and
commutative — so shard-parallel reductions combine to the same result as
one pass. `mean()`/`variance()` form n*sum_sq - sum^2 exactly in u128 when
it fits (always, at Hamming scale) and only round in the final division.

`Cascade::observe_batch(&[u32])` folds a whole batch into the rolling
floor with the Chan-Golub-LeVeque parallel-variance merge: same mu /
sigma / observations as per-element `observe` up to f64 rounding, and
shards may be folded in any order. Drift is judged once per batch (μ moved
more than 2σ of the pre-batch state), the test `observe` applies per step.

Tests: exact vs a u128 reference across chunk boundaries and full-range
values, u32::MAX worst case (100 003 squares of (2^32-1)^2), merge
exactness at every split in both orders, mean/variance vs two-pass f64;
observe_batch vs sequential observe (cold and warm), shard order
independence, and the alert firing on a shifted batch but not a stable one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
Operator rule: anything cosine-shaped pays the Fisher-Z entry tax before
it is thresholded, averaged, banded or ranked by margin — no exceptions,
lab and calibration included; anything bitpacked goes through HDR popcount
stacking on the Hamming distance itself. ndarray had no Fisher transform at
all; this adds the one canonical entry point.

- fisher_z / fisher_z_inv / hyperbolic_depth: the helix convention
  (`helix::fisher_z::Similarity`, mirrored by `jc::stats::fisher_2z`):
  ln form, r clamped to [-1+1e-9, 1-1e-9], NaN propagates.
- fisher_z_f32_batch: 16 lanes through F32x16 + simd_ln_f32, bit-identical
  to the scalar f32 formula. The f32 rim is 1 - 2^-24 (FISHER_CLAMP_F32):
  1 - 1e-9 rounds to exactly 1.0 in f32, where ln(1 - r) is -inf.
- hamming_null_z: (d - n/2) / (sqrt(n)/2) against Binomial(n, 1/2) — the
  random-vector null needs no calibration samples.

All re-exported from ndarray::simd. Tests pin atanh(0.5) and the 1e-9 rim
value, check oddness, monotonicity, finiteness past the rim, NaN
propagation, the inverse, batch/scalar bit identity across the 16-lane
boundary, and the binomial null at 0, -1 and +3 sigma.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
pearson / spearman / cosine stay raw for display; pearson_z / spearman_z /
cosine_z are the forms any threshold, average or comparison must use, and
the doc example now thresholds in z-space. Found by the cosine-tax census:
this was the one ndarray site consuming a raw cosine-shaped coefficient.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
@coderabbitai

coderabbitai Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Essentials

Run ID: bcb9bbea-6691-4844-9cb0-d4ff773a4197

📥 Commits

Reviewing files that changed from the base of the PR and between 2c91538 and aa3dfd7.

📒 Files selected for processing (7)
  • crates/wasm-simd-parity/src/lib.rs
  • src/hpc/cascade.rs
  • src/hpc/mod.rs
  • src/hpc/reliability.rs
  • src/hpc/statistics.rs
  • src/hpc/zspace.rs
  • src/simd.rs
 ______________________________________________________
< Weeks of programming can save you hours of planning. >
 ------------------------------------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

…ight to i8 codes

Remove fisher_z_f32_batch (wrote a z buffer) and fisher_z_inv (the way back
to cosine). Add ZGamma { z_min, z_range }: 8 LE bytes, the FamilyGamma
layout and code formula. fit() scans only the cosine min/max (Fisher-Z is
monotone), encode_batch() fuses clamp + transform + normalize in F32x16 and
writes only the i8 code, and threshold_code() moves a threshold into code
space once. fisher_z stays for single scalars.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
@AdaWorldAPI AdaWorldAPI changed the title zspace: Fisher-Z entry point + binomial Hamming null; batch Welford (moments_u32) recovered from #323 zspace: Fisher-Z entry tax without materializing z (ZGamma codes) + binomial Hamming null; batch Welford recovered from #323 Sep 23, 2026
@cursor

cursor Bot commented Sep 23, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_15974d73-f3de-4803-a11d-f88b772c16b3)

The FamilyGamma code inherited from bgz-tensor truncated toward zero, which
doubles the code-0 bucket and pulls every code half a step inward (measured
+0.49 / -0.49 mean error either side of centre), used -128 as a data level
reachable only by saturating below the envelope, and mapped NaN to 0, a real
code. Now: round-ties-even, saturate to -127..=127, NaN -> ZGamma::NAN_CODE
(-128). One shared quantize() keeps scalar and batch identical. The
FamilyGamma formula-parity test is replaced by a median-bias falsifier and a
NaN-sentinel test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
simd_ln_f32 is per-lane libm .ln(), so the Fisher path inherited the
platform libm: measured 624 of 8539 grid values (f32) and 566 (f64) differ by
1-2 ulp between x86-64 glibc and wasm32. Codes happened to agree on the grid,
but a one-ulp difference at a rounding boundary flips a code.

ln_det: musl/FreeBSD logf from IEEE add/sub/mul/div + bit ops in fixed order,
< 1 ulp over [2^-24, 2), used by both the scalar and batch encoders. Its x86
z digest now equals the wasm (musl-derived) libm digest. Codes and z digests
are pinned as constants. fisher_z (f64 scalar) keeps libm and is documented
as not bit-exact.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
CI's wasm job never runs cargo test, so the zspace golden digest was not
enforced on wasm32. selfcheck() now rebuilds the same 8539-value grid,
encodes it and compares against GOLDEN_CODES (rc 0x501), plus batch == scalar
per element (0x502).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
@AdaWorldAPI
AdaWorldAPI marked this pull request as ready for review September 24, 2026 06:07
@cursor

cursor Bot commented Sep 24, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_00acef45-6d8f-44ca-a39b-f70902f9b659)

@AdaWorldAPI
AdaWorldAPI merged commit e378938 into master Sep 24, 2026
26 of 27 checks passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: aa3dfd7365

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/hpc/cascade.rs
}
self.observations = old_n + b.n as usize;

if old_n > 10 && old_sigma > 0.0 && (self.mu - old_mu).abs() > 2.0 * old_sigma {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve alerts when a batch crosses the warm-up boundary

When a batch starts with 1–10 prior observations and brings the total above 10, checking old_n suppresses the shift alert even though observe checks the incremented observation count. For example, a cascade calibrated with 10 non-constant distances followed by one extreme distance can alert through observe but not observe_batch(&[distance]), making alert behavior depend on batching. Apply the warm-up check to the merged observation count instead.

Useful? React with 👍 / 👎.

AdaWorldAPI pushed a commit that referenced this pull request Sep 24, 2026
…polyfill-gni3cw"

This reverts commit e378938, reversing
changes made to 2c91538.

Restores the tree of 2c91538 exactly: removes hpc::zspace (ZGamma, ln_det,
fisher_z, hamming_null_z), the FidelityReport z accessors, moments_u32 /
MomentsU32 + Cascade::observe_batch, their simd.rs re-exports, and the
wasm-simd-parity ZGamma golden check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
AdaWorldAPI added a commit that referenced this pull request Sep 24, 2026
…-gni3cw

Revert #325 (zspace / ZGamma, FidelityReport z accessors, moments_u32 / observe_batch)
AdaWorldAPI pushed a commit that referenced this pull request Sep 24, 2026
…325

Mechanical recovery of the HDR/popcount hunks that #326 reverted, applied
unchanged:
- src/hpc/statistics.rs: MomentsU32 (exact integer n / Σx / Σx², order-
  independent merge, exact u128 variance) and moments_u32 (U64x8 lanes,
  split 32-bit squares, 2^28-chunk drain) with their tests.
- src/hpc/cascade.rs: Cascade::observe_batch (Chan-Golub-LeVeque merge of
  batch moments into the rolling floor, per-batch drift alert) with its
  three tests.
- src/simd.rs: the moments_u32 / MomentsU32 re-export line only.

Everything else from #325 stays reverted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
AdaWorldAPI pushed a commit that referenced this pull request Sep 24, 2026
Two defects in the code recovered from #325, both found in review:

- observe_batch gated its drift alert on `old_n > 10`, while observe
  counts the new element first and tests `> 10`, i.e. >= 10 prior. At
  exactly 10 prior observations, observe_batch(&[x]) stayed silent where
  observe(x) alerted. Gate is now `old_n >= 10`.
- MomentsU32::variance's fallback (when n*sum_sq or sum^2 overflows
  u128) subtracted two ~2^64-scale f64 values and could cancel the
  variance to 0 (2^33 values split between u32::MAX and u32::MAX-1:
  0.0 instead of 0.25). It now centres on q = floor(sum/n) exactly in
  u128 first, so the float step works on variance-sized quantities.

Each fix has a regression test that failed before the change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
AdaWorldAPI added a commit that referenced this pull request Sep 24, 2026
…-gni3cw

hdr: recover moments_u32 / MomentsU32 and Cascade::observe_batch (HDR-only harvest from #325)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants