Repository navigation
Conversation
Signed-off-by: JiaxinD <djx2048@gmail.com>
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. 🧰 Additional context used📚 Code guidelines (1)No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (4)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. 📝 Walkthrough
WalkthroughThe Llama changes add offline checkpoint-cache tests, release temporary tensors during checkpoint loading, and publish the prefill plan before compiling the decode plan. ChangesLlama checkpoint and build resource handling
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Refactor Suggested reviewers: Merge Risk: ⚪ Minimal · up to The reviewed changes have no established merge-blocking issue. The reported tests and checks passed, though real-checkpoint TensorRT compilation was not run for this head. 🚥 Pre-merge checks | ✅ 8 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (8 passed)
Comment |
Signed-off-by: JiaxinD <djx2048@gmail.com>
…etime Signed-off-by: JiaxinD <djx2048@gmail.com>
Signed-off-by: JiaxinD <djx2048@gmail.com>
Reuse the checkpoint-cache prerequisite from NVIDIA#1424 so Dev CI can validate the host-memory changes without Hub metadata access. Preserve exact revisions and config checks. Signed-off-by: JiaxinD <djx2048@gmail.com>
Signed-off-by: JiaxinD <djx2048@gmail.com>
Background
Llama compilation retains avoidable host allocations: FP32 embedding/projection sources during layer loading and serialized prefill bytes during decode compilation. These lifetimes create memory pressure without contributing to engine computation.
Exit Criteria
Release each temporary source before the next load and stage/release prefill bytes before decode compilation. Preserve mapped weights, named engine payloads and atomic publication. No GPU performance or whole-process RSS claim.
Implementation
prefill.planbefore decode compilation and release its in-memory bytes.41c5552c, applying the lifetime change inside the new native-builder callback while retaining complete Edge/paired dispatch. Keep the upstream mode-aware offline checkpoint resolver and our cache regressions.families/llama. Physical bundle section order changes, but names, payloads, reading by offset, ABI and bundle format remain unchanged.Change categories
Validation
Commands and Results
python -m pytest families/llama/tests/test_support.py families/llama/tests/test_checkpoint_cache.py core/builder/tests/test_bundle_writer.py core/builder/tests/test_bundle_reader.py tools/tests/test_architecture.py tools/tests/test_family_impact.py -q -p no:cacheproviderpython -m tools.model_ci validate,python -m tools.test_impact --validate, focused Ruff andgit diff --checkExisting regressions exercise the real BundleWriter/safetensors paths with compilation stubs: release before decode, unchanged distinct weights/overrides/tied embeddings, and retention of the published bundle after decode failure. The lifetime and cache regressions previously failed before their fixes; this run checks integration with the new upstream builder.
Earlier synthetic CPU
tracemallocexperiments reduced the two-plan peak from 33567185 to 16789012 bytes and the projection fixture from 16543495 to 10769185 bytes, with identical mapped-output SHA256. Those historical allocation experiments are not real-model RSS/GPU benchmarks for this merged head.Hardware, Environment, and Revisions
Head
2df0010a87f39b56496349256a430553bcf73ef2, tested tree4726bb5216dac054b29513d20d5f0116cecbfb90, based on upstream41c5552ce998f6d46a083e15da105258a78778c7. Linux/WSL, Python 3.12, CPU only. Commands usePYTHONPATH=core/builder:apps/benchmark:.. No model weights or datasets downloaded.Not Run / Remaining Gaps
No local real-checkpoint TensorRT compilation, GPU parity, throughput/latency or whole-process RSS measurement. Native Edge/speculative execution was retained from upstream, not newly hardware-qualified here. Previous-head Stable/Dev passed; new-head CI remains separate and must be checked. Radix authentication currently blocks GPU work.
Contributor Self-Review
Notes For Future Readers
Converted outputs remain owned by the weight dictionary. BundleWriter stages synchronously and its existing caller aborts temporary sections on failure. The synthetic tests verify numerical conversion and release timing; they do not establish GPU resource/performance equivalence. Scoped implementation and applicable CPU regression checks are complete; Draft remains due to the repository's unchanged Ready permission/capacity restriction.
Risk level
Allocation lifetimes and physical section order change. Numerical conversions, engine graphs and names remain unchanged, with focused output/publication coverage.