Describe cross-component tied weights - #747
Conversation
Expose canonical and alias parameter endpoints through component inspection so optimizers can coordinate independent builds. Materialize Gemma4 packed token embeddings into the split decoder LM head during ONNX export. Signed-off-by: Xiaoyu Zhang <xiaoyuzhang@microsoft.com>
Performance Comparison
|
🏗️ Architecture Diff
No architecture changes detected. ✅ Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed) |
Use Hugging Face tie_word_embeddings metadata together with explicit component source aliases to derive standard embedding/head sharing. Remove the Gemma4-specific shared-weight resolver while retaining its split export materialization. Signed-off-by: Xiaoyu Zhang <xiaoyuzhang@microsoft.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Critical typing and standalone component tie-handling issues remain unresolved, with additional unified-model test coverage needed.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
Open (1)
What changed in this PR
Adds cross-component shared-weight metadata and Gemma4 tied embedding/LM-head materialization for split exports.
Changes:
- Adds shared-weight metadata and validation to
inspect_components(). - Declares Gemma4 tied embeddings from configuration.
- Materializes packed embedding sidecars for decoder LM heads.
- Adds targeted inspection and quantized-weight tests.
| File | Summary |
|---|---|
src/mobius/models/gemma4.py |
Gemma4 shared-weight declarations and LM-head materialization |
src/mobius/models/gemma4_test.py |
Quantized tied-weight regression coverage |
src/mobius/_inspect.py |
Shared-weight metadata models and validation |
src/mobius/_inspect_test.py |
Inspection contract tests |
src/mobius/__init__.py |
Public API exports |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Carry inferred shared-weight groups on ComponentManifest and expand packed aliases before architecture-specific preprocessing. Remove Gemma4-specific materialization so standard tied component models use the common loading path. Signed-off-by: Xiaoyu Zhang <xiaoyuzhang@microsoft.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Quantization-only tied-weight declarations are not detected, preventing packed LM-head sidecar materialization.
Get a fresh assessment by requesting another Copilot review.
Review effort: Balanced
Findings: 1
Open (1)
Resolved since last review (1)
|
Cross-repository review of this PR together with microsoft/Olive#2684 (reviewed head
Cross-PR contract: Mobius recognizes an output module named This was a static review of the two PR diffs and immediate consumers, not an end-to-end Olive checkpoint -> Mobius ONNX runtime test. I found no confirmed blocking defect in this PR's tested Gemma4 path; the cross-build failure cases are posted on the Olive PR. |
Inspect both mapping and object quantization declarations, including parsed configuration, so a model-level false flag cannot hide a tied packed embedding/head relationship. Signed-off-by: Xiaoyu Zhang <xiaoyuzhang@microsoft.com>
|
Follow-up review at head The new tests cover ties recorded only in quantization metadata, including the important case where the model-level flag is false but the packed checkpoint records a tie. However, The endpoint inference still uses explicit I have not run a combined Olive assembled-checkpoint -> Mobius ONNX export/runtime validation. This is a follow-up comment, not a formal GitHub changes-requested review. |
Keep explicit quantized tying authoritative, but prefer the nested text flag over the parent flag for unquantized models. State that cross-component inference requires unique explicit source aliases and cover Gemma4 and Gemma4-unified conflicts against config parsing. Signed-off-by: Xiaoyu Zhang <xiaoyuzhang@microsoft.com>
|
@titaiwangms Addressed in fb6b087: for unquantized sources the nested text flag now takes precedence over the parent (matching Gemma4Config parsing); an explicit quantization-level tie still restores sharing. Tests cover conflicting flags for both Gemma4 variants. Inference is scoped to models exposing a unique explicit embedding/head source-path alias pair, now stated in the public inspection contract. The companion microsoft/Olive#2684 includes a real tiny-Gemma4 Olive checkpoint -> Mobius paired ONNX -> ORT CPU full-logit parity test. |
titaiwangms
left a comment
There was a problem hiding this comment.
Reviewed the latest tied-weight inference and explicit alias scope. The earlier tie-precedence concern is resolved, and the PR checks are green. Approved.
## Summary - preserve one Hugging Face component per build while deferring aliases only when the canonical table is planned for compatible quantization; reject incompatible layouts before running builds - automatically select scoped LM-head, embedding, and vision targets for KQuant/RTN while retaining explicit opt-outs and whole-model defaults - fail if required shared-weight assembly is skipped, keep deferred aliases through follow-up passes, recover missing float heads for one-sided quantization, and validate shared-weight ownership and metadata - recreate packed secondary embedding placeholders on Hugging Face reload when their sidecars are present ## Validation - targeted quantizer, assembly, secondary-embedding and cross-repository suites (187 passing); Ruff and lintrunner PYLINT - network-free tiny Gemma4: actual Olive two-build assembly -> Mobius paired embedding/decoder export -> ONNX Runtime CPU inference; full logits agree with reloaded Hugging Face model (maximum absolute difference ~4.2e-7) Depends on onnxruntime/mobius#747 for the shared-weight contract and generic export materialization. --------- Signed-off-by: Xiaoyu Zhang <xiaoyuzhang@microsoft.com> Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>


Summary
inspect_components()Validation
Companion: microsoft/Olive#2684 handles component quantization and assembly.