Repository navigation
fix(aiperf): report benchmark timings directly - #1630
Open
chaofengw-nv wants to merge 1 commit into
Open
chaofengw-nv wants to merge 1 commit into
chaofengw-nv wants to merge 1 commit into
Conversation
Report Native and TRTMC benchmark p50 as measurements while retaining work differences and partial coverage as diagnostics. Apply accuracy capacity exclusions to both timing sides and refresh saved reports without inference or changes to accuracy gates. Signed-off-by: chaofengw <chaofengw@nvidia.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Background
Benchmark qualification already uses one AIPerf response stream for quality and task-call timings. Its report still treats unequal generation lengths as an invalid performance result and treats capacity exclusions accepted by accuracy as missing timing responses. This contradicts the requested report: benchmark accuracy plus Native/TRTMC p50 from the same execution.
Exit Criteria
Implementation
Change categories
Validation
Commands and Results
PYTHONPATH=/tmp/trtmc-aiperf-accuracy-venv/lib/python3.12/site-packages:/home/chaofengw/workspace/TensorRT-Model-Connect/.venv-trtmc/lib/python3.12/site-packages:apps/aiperf_qual:apps/aiperf_qual/plugins /tmp/trtmc-aiperf-accuracy-venv/bin/python -m pytest apps/aiperf_qual/tests -q: 308 passed, 1 skipped.ruff check --config ruff.toml apps/aiperf_qual/trtmc_aiperf_qual/{accuracy_recovery,benchmark_perf,cli,execution,judge,report,report_html,runner}.py apps/aiperf_qual/tests/{test_execution,test_benchmark_perf}.py: passed.git diff --check: passed.Hardware, Environment, and Revisions
CPU validation uses Python 3.12 and the prepared AIPerf 0.13.0 qualification environment. The branch starts from
54d77286e(#1628); tested implementation head is71a78cd93. Offline replay uses original GB300 benchmark records captured by source revision587840c63with the conversation-identity fix applied; no new GPU benchmark is claimed.Not Run / Remaining Gaps
Changes are not deployed to the active GB300 campaigns. Full campaign inference continues with its current code. Fixed-workload diagnostics remain outside this benchmark-report correction. The COCO accuracy dependency test is skipped because
pycocotoolsis absent from the CPU environment.Contributor Self-Review
Notes For Future Readers
Review the execution scope and measurement verdict before the refresh/reporting paths. Rejudge without an environment intentionally retains the recorded accuracy contract. Reports without execution evidence may reapply their saved aggregate verdict, but capacity reselection requires original execution and raw rejection records. No models, checkpoints, or bundles need rebuilding for this change.
Risk level
The observable Perf status and the Native timing sample population change when accuracy excluded capacity rejections. Rejection counts, attempted counts, and work differences remain visible; missing or inconsistent capacity evidence fails rather than guessing a sample selection.