Skip to content

test(llama): add pinned MiniCPM5-1B coverage - #1424

Draft
JiaxinD wants to merge 8 commits into
NVIDIA:mainfrom
JiaxinD:test/llama-minicpm5-1b
Draft

JiaxinD wants to merge 8 commits into
NVIDIA:mainfrom
JiaxinD:test/llama-minicpm5-1b

Conversation

@JiaxinD

@JiaxinD JiaxinD commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Background

Llama has MiniCPM5-2B coverage but lacks a pinned 1B case. MiniCPM5-1B has hidden width 1536 and attention width 2048; its tokenizer also exposed a prior 18-versus-16-token prefill mismatch.

Exit Criteria

  • Select both pinned 1B cases through the existing harness.
  • Preserve the HF comparison and raw prompt's 16-token native prefill assertion.
  • Build and pass the selected cases on target GPU hardware.

Implementation

  • Add FP16 manifests with FP32 references, publisher config fixtures, CPU geometry tests and an explicit release-performance exclusion.
  • Add a raw multilingual/digit prompt with fixed HF IDs and a native prefill-count receipt.
  • Include the original fix(llama): scan all Split steps for pretokenizer variant classification #1423 tokenizer prerequisite, retaining its author and published history.
  • Merge upstream 0fec6d03, preserving speculative/Edge execution and the Qwen embedding exclusion. Keep mode-aware offline snapshot resolution and cache regressions.
  • No shared runtime/API/ABI or bundle-format change. The performance exclusion is the only change outside families/llama.

Change categories

  • CI or developer tooling
  • Model or runtime behavior

Validation

Commands and Results

Command Result
python -m pytest families/llama/tests/test_checkpoint_cache.py families/llama/tests/test_minicpm5_config.py families/llama/tests/test_runtime_receipt.py qualification_tests/benchmark_qualification/performance/tests -q -p no:cacheprovider --junitxml=minicpm-cpu.xml 570 passed
cmake --build build-sm90 --parallel 8 --target trtmc trtmc_backend_trt trtmc_model_llama test_llama_pretokenizer_variant test_llama_native_kv_cache test_llama_chat_template test_llama_plugin_helpers test_llama_pipeline test_llama_speculative_contract Passed
ctest --test-dir build-sm90 --output-on-failure -R '^test_llama_(pretokenizer_variant|native_kv_cache|chat_template|plugin_helpers|pipeline|speculative_contract)$' 6 passed, 0 skipped
python -m pytest families/llama/tests/test_e2e.py --e2e-model minicpm5-1b -vv -s --basetemp=/work/e2e-temp --junitxml=/work/evidence/e2e.xml 2 selected cases passed, 8 unrelated cases skipped
python -m tools.model_ci validate, python -m tools.test_impact --validate, git diff --check Passed

The raw case reports [trtmc.prefill] tokens=16 launches=1 max_chunk=16; all 10 generated IDs match its FP32 HF reference. JUnit, case receipts, source/binary hashes and environment records are retained. No thresholds were changed.

Hardware, Environment, and Revisions

Head d0d47785b92fe881f2a34cb16052396b86e56e32, tree 3c305f19fa5ff6a7b89fc6ef57bb05538649db57. openbmb/MiniCPM5-1B revision 87179e5c1f455ef22e6223592d2d61351b525bfc, Apache-2.0, staged offline. One H200/SM90, Ubuntu 24.04, Python 3.12, driver 595.71.05, CUDA 13.3, TensorRT 11.1.0.106, PyTorch 2.12.0+cu130, Transformers 5.2.0. Image nvcr.io/nvidia/tensorrt:26.07-py3@sha256:b82db1abc23750ab0069abc99bbe4ea29138dbdc23ea39861199e2346638b48a. GPU pinned to the assigned UUID and visible index 0.

Not Run / Remaining Gaps

Long-context, alternate precision, other checkpoints and performance qualification were not run. Current public/protected CI are separate from this local target-hardware result.

Contributor Self-Review

  • I have completed a self-review of this change.

Notes For Future Readers

Review #1423 first. The receipt proves prefill count, not every native input ID; the HF generation comparison remains independent. Scoped GPU validation is complete; repository review capacity currently keeps this PR draft.

Risk level

  • Medium

Tokenizer behavior and new selected cases affect this checkpoint; broader model/runtime qualification is outside scope.

roma5087 and others added 2 commits September 25, 2026 08:37
detect_split_variant() returned on the first Split step found in a
Sequence pre_tokenizer, even when that step's regex classified as the
generic kLlama default. Checkpoints that isolate digit-grouping
(\p{N}{1,3}) as its own Split step ahead of the step whose regex
actually identifies the variant (e.g. Qwen3's [^\r\n... signature)
were misclassified as kLlama, silently disabling variant-specific
pre-tokenization behavior for the whole checkpoint.

Confirmed directly against openbmb/MiniCPM5-2B's real, published
tokenizer.json, which has exactly this shape.

Now scans every Split step for classification instead of only the
first, and extracts digit-group size independently of which step
decides the variant, since the two are not guaranteed to be the same
step (confirmed: they aren't, in MiniCPM5-2B's real tokenizer.json).
Throws if two different Split steps would classify to two different
non-kLlama variants, matching this file's existing convention of
throwing on ambiguous/unsupported pre-tokenizer contracts.

New test (test_llama_pretokenizer_variant.cpp, CPU-only, no GPU/TRT
dependency) uses MiniCPM5-2B's real pre_tokenizer field verbatim to
verify both the newline-attachment and digit-grouping behavior this
fix restores, plus a synthetic fixture for the ambiguous-variant
throw case. Mutation-checked against the original code (all three
checks fail on unpatched bpe_tokenizer.cpp, pass on the fix).

Scoped to families/llama only; the same detect_split_variant/
classify_split pattern is duplicated per-family across most other
families' own bpe_tokenizer.cpp copies, per this repo's own
per-family duplication convention. Fixing those is out of scope here.

Signed-off-by: Matthew Romano <matthewromano5087@gmail.com>
Signed-off-by: JiaxinD <djx2048@gmail.com>
@coderabbitai

coderabbitai Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Signed-off-by: JiaxinD <djx2048@gmail.com>
Signed-off-by: JiaxinD <djx2048@gmail.com>
Merge the original NVIDIA#1423 commit for combined MiniCPM5-1B qualification. Preserve its author and commit identity; retain the native prompt-count regression and unchanged generation thresholds.

Signed-off-by: JiaxinD <djx2048@gmail.com>
JiaxinD added a commit to JiaxinD/TensorRT-Model-Connect that referenced this pull request Sep 29, 2026
Reuse the checkpoint-cache prerequisite from NVIDIA#1424 so Dev CI can validate the host-memory changes without Hub metadata access. Preserve exact revisions and config checks.

Signed-off-by: JiaxinD <djx2048@gmail.com>
Signed-off-by: JiaxinD <djx2048@gmail.com>

# Conflicts:
#	qualification_tests/benchmark_qualification/performance/config/release.yaml
Signed-off-by: JiaxinD <djx2048@gmail.com>
Signed-off-by: JiaxinD <djx2048@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants