Fix Day 1 GPU float32 attention test under default MLX precision - #324
Merged
Merged
Conversation
Document the fresh-checkout learner and reference build prerequisites in setup, the Day 1 lesson, and the preface. Keep the starter TODO boundary explicit. AI-Assisted: GPT-6 Sol + Sentinel
Move the existing setup prerequisite above the first reference test example so the preface is runnable in reading order. AI-Assisted: GPT-6 Sol + Sentinel
skyzh
marked this pull request as ready for review
September 27, 2026 01:03
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
2**-11absolute tolerance under MLX's default precision mode. This empirical margin covers the reported1.1e-4gap for the bounded fixture; it is not a hardware or general attention error bound. Scale and mask cases, CPU and float16 cases, and the shared comparator retain their strict checks.testandtest-refsolPDM scripts use MLX's default precision mode; neither setsMLX_ENABLE_TF32. Learner and reference implementations are unchanged.Related to #322. This PR does not close the issue: the reported Apple M5 / MLX 0.32.1 case still needs a direct retest.
Evidence
MLX_ENABLE_TF32unset, focused Day 1 reference tests pass 36/36. The full unfiltered reference suite on the executable-equivalent parent passes 679 with 2 optional skips. The unchanged learner starter collects, with 4 softmax passes and 32 expected attention TODO failures; a temporary implementation of the lesson formula passes all 36 Task 1 cases.+1.1e-4offset and rejects+1e-3; scale and mask checks reject the smaller offset.b6c8ea5465e1347d02827e16a3c523c6b5e78c3e. Hosted PR checks and public-head review are separate gates.Limits
No direct Apple M5 / MLX 0.32.1 retest was available. The local synthetic rounding surrogate is diagnostic evidence, not a bound on MLX or M5. This PR makes no BF16 or model-performance claim.