Skip to content

Fix Day 1 GPU float32 attention test under default MLX precision - #324

Merged
skyzh merged 5 commits into
mainfrom
atlas/week1-day1-f32-attention-20260925
Sep 27, 2026
Merged

skyzh merged 5 commits into
mainfrom
atlas/week1-day1-f32-attention-20260925

Conversation

@skyzh

@skyzh skyzh commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Compare the Week 1 Day 1 Task 1 unmasked float32 GPU attention output with a scoped 2**-11 absolute tolerance under MLX's default precision mode. This empirical margin covers the reported 1.1e-4 gap for the bounded fixture; it is not a hardware or general attention error bound. Scale and mask cases, CPU and float16 cases, and the shared comparator retain their strict checks.
  • The test and test-refsol PDM scripts use MLX's default precision mode; neither sets MLX_ENABLE_TF32. Learner and reference implementations are unchanged.
  • Keep the setup, preface, and Day 1 instructions that build both native extensions before the first tests.

Related to #322. This PR does not close the issue: the reported Apple M5 / MLX 0.32.1 case still needs a direct retest.

Evidence

  • On M4 Pro / MLX 0.32.0 with MLX_ENABLE_TF32 unset, focused Day 1 reference tests pass 36/36. The full unfiltered reference suite on the executable-equivalent parent passes 679 with 2 optional skips. The unchanged learner starter collects, with 4 softmax passes and 32 expected attention TODO failures; a temporary implementation of the lesson formula passes all 36 Task 1 cases.
  • Deliberately wrong scale, mask, batch, shape, and NaN outputs are rejected by the affected comparison. The scoped plain comparison accepts a +1.1e-4 offset and rejects +1e-3; scale and mask checks reject the smaller offset.
  • Independent exact-local-head numerical, factual, learner, and static reviews passed on b6c8ea5465e1347d02827e16a3c523c6b5e78c3e. Hosted PR checks and public-head review are separate gates.

Limits

No direct Apple M5 / MLX 0.32.1 retest was available. The local synthetic rounding surrogate is diagnostic evidence, not a bound on MLX or M5. This PR makes no BF16 or model-performance claim.

Document the fresh-checkout learner and reference build prerequisites in setup, the Day 1 lesson, and the preface. Keep the starter TODO boundary explicit.

AI-Assisted: GPT-6 Sol + Sentinel
Move the existing setup prerequisite above the first reference test example so the preface is runnable in reading order.

AI-Assisted: GPT-6 Sol + Sentinel
@skyzh skyzh changed the title Make Day 1 float32 attention tests use full MLX precision Fix Day 1 GPU float32 attention test under default MLX precision Sep 26, 2026
@skyzh
skyzh marked this pull request as ready for review September 27, 2026 01:03
@skyzh
skyzh merged commit da3e841 into main Sep 27, 2026
1 check passed
@skyzh
skyzh deleted the atlas/week1-day1-f32-attention-20260925 branch September 27, 2026 01:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant