zspace: Fisher-Z entry tax without materializing z (ZGamma codes) + binomial Hamming null; batch Welford recovered from #323 - #325
Conversation
`moments_u32(&[u32]) -> MomentsU32 { n, sum, sum_sq }`: exact integer
moments, eight lanes at a time through U64x8. Squares are split into
32-bit halves so every lane add stays below 2^32; lanes drain into u128
totals every 2^28 chunks, before an 8-lane reduce_sum could overflow u64.
`MomentsU32::merge` is plain integer addition — exact, associative and
commutative — so shard-parallel reductions combine to the same result as
one pass. `mean()`/`variance()` form n*sum_sq - sum^2 exactly in u128 when
it fits (always, at Hamming scale) and only round in the final division.
`Cascade::observe_batch(&[u32])` folds a whole batch into the rolling
floor with the Chan-Golub-LeVeque parallel-variance merge: same mu /
sigma / observations as per-element `observe` up to f64 rounding, and
shards may be folded in any order. Drift is judged once per batch (μ moved
more than 2σ of the pre-batch state), the test `observe` applies per step.
Tests: exact vs a u128 reference across chunk boundaries and full-range
values, u32::MAX worst case (100 003 squares of (2^32-1)^2), merge
exactness at every split in both orders, mean/variance vs two-pass f64;
observe_batch vs sequential observe (cold and warm), shard order
independence, and the alert firing on a shifted batch but not a stable one.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
Operator rule: anything cosine-shaped pays the Fisher-Z entry tax before it is thresholded, averaged, banded or ranked by margin — no exceptions, lab and calibration included; anything bitpacked goes through HDR popcount stacking on the Hamming distance itself. ndarray had no Fisher transform at all; this adds the one canonical entry point. - fisher_z / fisher_z_inv / hyperbolic_depth: the helix convention (`helix::fisher_z::Similarity`, mirrored by `jc::stats::fisher_2z`): ln form, r clamped to [-1+1e-9, 1-1e-9], NaN propagates. - fisher_z_f32_batch: 16 lanes through F32x16 + simd_ln_f32, bit-identical to the scalar f32 formula. The f32 rim is 1 - 2^-24 (FISHER_CLAMP_F32): 1 - 1e-9 rounds to exactly 1.0 in f32, where ln(1 - r) is -inf. - hamming_null_z: (d - n/2) / (sqrt(n)/2) against Binomial(n, 1/2) — the random-vector null needs no calibration samples. All re-exported from ndarray::simd. Tests pin atanh(0.5) and the 1e-9 rim value, check oddness, monotonicity, finiteness past the rim, NaN propagation, the inverse, batch/scalar bit identity across the 16-lane boundary, and the binomial null at 0, -1 and +3 sigma. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
pearson / spearman / cosine stay raw for display; pearson_z / spearman_z / cosine_z are the forms any threshold, average or comparison must use, and the doc example now thresholds in z-space. Found by the cosine-tax census: this was the one ndarray site consuming a raw cosine-shaped coefficient. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Essentials Run ID: 📒 Files selected for processing (7)
✨ Finishing Touches📝 Generate docstrings
Comment |
…ight to i8 codes
Remove fisher_z_f32_batch (wrote a z buffer) and fisher_z_inv (the way back
to cosine). Add ZGamma { z_min, z_range }: 8 LE bytes, the FamilyGamma
layout and code formula. fit() scans only the cosine min/max (Fisher-Z is
monotone), encode_batch() fuses clamp + transform + normalize in F32x16 and
writes only the i8 code, and threshold_code() moves a threshold into code
space once. fisher_z stays for single scalars.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_15974d73-f3de-4803-a11d-f88b772c16b3) |
The FamilyGamma code inherited from bgz-tensor truncated toward zero, which doubles the code-0 bucket and pulls every code half a step inward (measured +0.49 / -0.49 mean error either side of centre), used -128 as a data level reachable only by saturating below the envelope, and mapped NaN to 0, a real code. Now: round-ties-even, saturate to -127..=127, NaN -> ZGamma::NAN_CODE (-128). One shared quantize() keeps scalar and batch identical. The FamilyGamma formula-parity test is replaced by a median-bias falsifier and a NaN-sentinel test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
simd_ln_f32 is per-lane libm .ln(), so the Fisher path inherited the platform libm: measured 624 of 8539 grid values (f32) and 566 (f64) differ by 1-2 ulp between x86-64 glibc and wasm32. Codes happened to agree on the grid, but a one-ulp difference at a rounding boundary flips a code. ln_det: musl/FreeBSD logf from IEEE add/sub/mul/div + bit ops in fixed order, < 1 ulp over [2^-24, 2), used by both the scalar and batch encoders. Its x86 z digest now equals the wasm (musl-derived) libm digest. Codes and z digests are pinned as constants. fisher_z (f64 scalar) keeps libm and is documented as not bit-exact. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
CI's wasm job never runs cargo test, so the zspace golden digest was not enforced on wasm32. selfcheck() now rebuilds the same 8539-value grid, encodes it and compares against GOLDEN_CODES (rc 0x501), plus batch == scalar per element (0x502). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_00acef45-6d8f-44ca-a39b-f70902f9b659) |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: aa3dfd7365
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| } | ||
| self.observations = old_n + b.n as usize; | ||
|
|
||
| if old_n > 10 && old_sigma > 0.0 && (self.mu - old_mu).abs() > 2.0 * old_sigma { |
There was a problem hiding this comment.
Preserve alerts when a batch crosses the warm-up boundary
When a batch starts with 1–10 prior observations and brings the total above 10, checking old_n suppresses the shift alert even though observe checks the incremented observation count. For example, a cascade calibrated with 10 non-constant distances followed by one extreme distance can alert through observe but not observe_batch(&[distance]), making alert behavior depend on batching. Apply the warm-up check to the merged observation count instead.
Useful? React with 👍 / 👎.
…polyfill-gni3cw" This reverts commit e378938, reversing changes made to 2c91538. Restores the tree of 2c91538 exactly: removes hpc::zspace (ZGamma, ln_det, fisher_z, hamming_null_z), the FidelityReport z accessors, moments_u32 / MomentsU32 + Cascade::observe_batch, their simd.rs re-exports, and the wasm-simd-parity ZGamma golden check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
…-gni3cw Revert #325 (zspace / ZGamma, FidelityReport z accessors, moments_u32 / observe_batch)
…325 Mechanical recovery of the HDR/popcount hunks that #326 reverted, applied unchanged: - src/hpc/statistics.rs: MomentsU32 (exact integer n / Σx / Σx², order- independent merge, exact u128 variance) and moments_u32 (U64x8 lanes, split 32-bit squares, 2^28-chunk drain) with their tests. - src/hpc/cascade.rs: Cascade::observe_batch (Chan-Golub-LeVeque merge of batch moments into the rolling floor, per-batch drift alert) with its three tests. - src/simd.rs: the moments_u32 / MomentsU32 re-export line only. Everything else from #325 stays reverted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
Two defects in the code recovered from #325, both found in review: - observe_batch gated its drift alert on `old_n > 10`, while observe counts the new element first and tests `> 10`, i.e. >= 10 prior. At exactly 10 prior observations, observe_batch(&[x]) stayed silent where observe(x) alerted. Gate is now `old_n >= 10`. - MomentsU32::variance's fallback (when n*sum_sq or sum^2 overflows u128) subtracted two ~2^64-scale f64 values and could cancel the variance to 0 (2^33 values split between u32::MAX and u32::MAX-1: 0.0 instead of 0.25). It now centres on q = floor(sum/n) exactly in u128 first, so the float step works on variance-sized quantities. Each fix has a regression test that failed before the change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
…-gni3cw hdr: recover moments_u32 / MomentsU32 and Cascade::observe_batch (HDR-only harvest from #325)
Why
Operator rules:
Before this PR, ndarray had no Fisher transform at all.
This PR also carries two commits that were meant for #323 but missed it: #323 was merged at
edbd97f, one push before them. They are cherry-picked here unchanged:moments_u32+Cascade::observe_batch, and their re-export.What
hpc::zspace(7bb45fb→bb14de6→fd45c89→f2a461a)ZGamma { z_min, z_range }is the z-space envelope: 8 little-endian bytes.fit(cosines)scans only the cosine min and max. Fisher-Z is monotone, so no z value is ever stored.encode/encode_batchgo straight from cosine to ani8code, and only the code is written. Both call one sharedquantize().((z−z_min)/z_range)·254 − 127, rounded ties-to-even and saturated to the symmetric range−127..=127.−128is reserved asZGamma::NAN_CODE.threshold_code(c)converts a threshold into code space once. There is no decode back to cosine.Bit-exact across targets (
f2a461a).simd_ln_f32is a per-lane libm.ln(), so the first version inherited the platform libm. Measured over an 8,539-value grid:The codes happened to agree on that grid, but a one-ulp difference at a rounding boundary flips a code.
ln_detreplaces libm in the code path. It is musl/FreeBSDlogf, built only from IEEE add, sub, mul, div and bit operations in a fixed order. Rust never contracts these into FMA, so every conforming target gets the same bits. Its error is under 1 ulp over its domain[2⁻²⁴, 2). Its x86 z digest now equals the digest wasm's (musl-derived) libm produced before.Where this departs from bgz-tensor. It shares
FamilyGamma's byte layout, but not its code.FamilyGammapredates this analysis and was not calibrated for exactness:FamilyGammabehaviour−128used as a data levellnUnchanged:
fisher_z/hyperbolic_depthstay, for single f64 scalars only (a report statistic, a threshold). They keep libm and are documented as not bit-exact.hamming_null_zstays. It is the binomial-null z for bitpacked data.Removed:
fisher_z_f32_batchwrote a buffer of z values.fisher_z_invwas the way back to cosine.FidelityReportpays the tax (b8babbd)pearson_z/spearman_z/cosine_z(scalar statistics).Batch Welford, recovered (
90f1128,414aa8e)moments_u32computes exact integer(n, Σx, Σx²)throughU64x8.MomentsU32::mergeis exact and order-independent.Cascade::observe_batchuses the Chan–Golub–LeVeque merge.wasm enforcement (
aa3dfd7)CI never runs
cargo teston wasm. Socrates/wasm-simd-parity's node selfcheck now rebuilds the grid and checks the pinned codes digest (rc0x501) and batch == scalar (rc0x502).Tests
zspace(14 tests):ln_detis within 1 ulp across its domain, sampling every 257th bit pattern;GOLDEN_CODES,GOLDEN_Z32) that must hold on every target.Disable checks — each deliberate bug makes the listed test fail:
.ln()back in the scalar path0x501Test runs by build:
config-v3(AVX2)config-v4(AVX-512)The doc tests pass. clippy
-D warningsand fmt are clean.Open questions (not decided in this PR)
code = round(127·z/z_absmax): the sign of a code becomes the sign of the evidence. It costs half the resolution when a population is one-sided.fisher_z·√(n−3)andhamming_null_zare both ~N(0,1) under their nulls.simd_ln_f32itself is still per-lane libm, used by callers outside zspace. Whether the facade'slnshould becomeln_det(bit-exact, and vectorizable in the same op order) is a separate decision.Not in this PR
similarity_z/fisher_z_inverse.FamilyGamma: truncation,−128, NaN and libm are all present there too. Fixing it changes stored codes, so it needs its own decision.🤖 Generated with Claude Code
https://claude.ai/code/session_019HnekoM1EidTwQLS3oFVFm
Summary by CodeRabbit