fix(quantizer): support preserved MXFP4 in DSpark plans - #645
Open
apetersson wants to merge 1 commit into
Open
Conversation
Allow native packed MXFP4 only for routed DSpark expert tensors while keeping ordinary MXFP4 quantization requests rejected. Validate the paired I8 weight and F8_E8M0 scale contract before planning or conversion. Add native planner and repacker regressions plus mixed IQ2_XXS/MXFP4 Metal coverage for one through five tokens. Correct the fused tiny gate/up diagnostic slot indexing exercised by those tests. Fixes antirez#642
Author
|
Speed regression benchmark (M1 Ultra, 7 alternating base/patched DSpark pairs): |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Enable preserved native MXFP4 for DSpark routed experts only, with strict I8/F8_E8M0 shape, payload, and alignment validation. Non-expert MXFP4 remains rejected.
Add mixed IQ2_XXS gate/up + MXFP4 down Metal coverage for 1–5 tokens, harden block decoding fixtures, and fix fused tiny gate/up output indexing. No GGUF format changes; compatibility with
ds4f-mxfp4is preserved.Validation
git diff --checkpasses.Fixes #642