Skip to content

simd: mask_ternlog_popcount / mask_ternlog_any — Boolean membership ends in Count/Any with no mask written - #322

Merged
AdaWorldAPI merged 1 commit into
masterfrom
claude/fold-distillation-pr-wave-s57uj7
Sep 23, 2026
Merged

AdaWorldAPI merged 1 commit into
masterfrom
claude/fold-distillation-pr-wave-s57uj7

Conversation

@AdaWorldAPI

Copy link
Copy Markdown
Owner

What

Two slice functions in simd_masking_ops.rs, re-exported from ndarray::simd:

  • mask_ternlog_popcount::<IMM>(a, b, c) -> u64: Σ popcount(ternlog(a, b, c))
  • mask_ternlog_any::<IMM>(a, b, c) -> bool: is any bit of ternlog(a, b, c) set?

A 2- or 3-input Boolean membership can now end in Count or Any without writing a mask. Two-input functions are the same call with a table that ignores c.

The question it answers

Is a new T1 primitive needed for "Boolean membership → Count/Any", or can the existing U64x8 operations already do it without touching memory?

Measured answer: nothing new is needed below the slice loop. U64x8::{ternlog, popcnt, +, |, reduce_sum} already exist in every backend (avx512, avx2-polyfill, scalar, neon, wasm). These two functions are loops over those methods, so no backend code, cfg or intrinsics are added.

Contract

  • Word-level: the result is identical to mask_ternlog + popcount_batch_u64 / mask_any over the materialized mask.
  • Odd truth tables: an odd IMM sets the dead tail bits of the last word, and those bits are counted. The caller masks the last word, the same rule mask_ternlog already documents for its output.
  • Register padding is never counted. An odd table turns an all-zero padding lane into all-ones. The tail therefore runs through the packed op, and only the real lanes are read.
  • mask_ternlog_any checks once per block of 8 chunks, not once per chunk. A per-chunk check lost to the materializing pair on avx2 (0.91×).

Evidence

Probe: examples/ternlog_fold_probe.rs. Truth table AND2_OR; timings are medians; "M/R" is the speed-up over materializing the mask and then reducing it.

backend words Count M/R Any M/R
avx2 (config-v3) 16 384 1.27 3.24
avx2 262 144 1.50 6.18
avx512 (config-v4, no vpopcntdq) 16 384 1.90 2.59
avx512 262 144 1.94 4.39

A plain scalar fused loop roughly ties the register fold on avx2 and at the memory-bound size on avx512. The win comes from not writing the mask, not from SIMD as such.

Tests

  • Differential: mask_ternlog_folds_match_the_materializing_pair_for_all_256_tables compares each fold against the exact pair it replaces. It covers all 256 truth tables × 14 lengths (0 … 129) × dense and sparse inputs.
  • Padding: mask_ternlog_folds_never_count_register_padding covers NOR3 over 9 words (7 padding lanes), hits that sit only in the tail, and hits past the first block.
  • Length-mismatch panics for both functions.
  • Disable runs, each failing when broken as intended: counting padding lanes, testing padding lanes in the Any tail, and removing Any's block check.
  • Parity: crates/simd-masking-parity slice_ternlog! now checks both folds (codes 0x69x count, 0x65x any). It passes on native v4, native v3, neon-qemu, wasm and wasm-scalar.
  • clippy -D warnings passes on v3 and v4.

Consumer

lance-graph-mask-risc will use these for its fused terminal path MaskOp::{And, Or, Xor, AndNot, Ternlog} → Count/Any, in a follow-up lance-graph PR.

🤖 Generated with Claude Code

https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG

…mbership ends in Count/Any with no mask written

Built only from U64x8 methods every realization already carries (ternlog,
popcnt, +, |, reduce_sum), so no backend code is added: the missing piece was
the slice loop, not an ISA primitive. Word-level contract identical to
mask_ternlog + popcount_batch_u64 / mask_any; register padding is never
counted (an odd table evaluates padding lanes to all-ones).

Probe examples/ternlog_fold_probe.rs: Count 1.3-1.9x, Any 2.6-6.2x over the
materializing pair on avx2/avx512. Tests cover all 256 tables x 14 lengths x
dense/sparse against the pair they replace; parity crate extended, green on
v4, v3, neon-qemu, wasm, wasm-scalar.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GXUahz73MZxtxWcfpHp9dG
@coderabbitai

coderabbitai Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Essentials

Run ID: 37907a7e-14ea-4b05-a48a-8990ed069a59

📥 Commits

Reviewing files that changed from the base of the PR and between 58a9fd2 and 7695fa6.

📒 Files selected for processing (6)
  • .claude/blackboard.md
  • Cargo.toml
  • crates/simd-masking-parity/src/lib.rs
  • examples/ternlog_fold_probe.rs
  • src/simd.rs
  • src/simd_masking_ops.rs
 ________________________________
< My GPUs mine bugs, not Crypto. >
 --------------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@cursor

cursor Bot commented Sep 23, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_1a49c2b0-e342-441f-b798-0361c9373764)

@AdaWorldAPI
AdaWorldAPI marked this pull request as ready for review September 23, 2026 15:37
@AdaWorldAPI
AdaWorldAPI merged commit 6ae40cd into master Sep 23, 2026
27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants