Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
36 changes: 20 additions & 16 deletions apps/aiperf_qual/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,9 @@ It does not use `qualification_tests/benchmark_qualification`.
(`noninferiority.py`): `pass` when TRTMC's regression is shown to be below the benchmark's margin, `fail`
when it is shown to exceed it, `inconclusive` otherwise. Tasks without a gold set compare outputs with
the native model (conversion parity). Random-weight test models are Perf only (`accuracy_source: none`).
- **Perf**: TRTMC must be faster than the native model (eager) at the candidate's precision: the speedup's
90% interval lies above 1.05 x (1 + guard) (the 5% margin widened by the largest server-instance and order effect
the order check measured, `guard_percent`), on every timed request, with the same work on both sides.
- **Perf**: report Native and TRTMC task-call p50 from the same benchmark responses used for Acc,
at the recorded effective precision. Timing is a measurement, not an additional acceptance gate.
Output-length differences, execution conditions, and partial coverage remain explicit observations.

## Design

Expand Down Expand Up @@ -69,9 +69,9 @@ A suite with `base: catalog` overrides the profile's catalog request with its da
- Every AIPerf run has a deadline: three times the profile's seconds in the run's ledger (`run-all --ledger`, at
least ten minutes), else 12 hours; a GPU phase that fails before producing its result runs once more.
- `summary` reports one result per model, worst first: White (no verdict: an error or a failed build; or no valid
comparison: the native model below a benchmark's floor, a Task without an Acc check, timings that cannot be
compared), Red (Acc or Perf worse than native beyond its margin), Yellow (Perf about equal to native, which counts
as a pass, or an Acc difference not shown either way), Green (quality passes with valid comparable timings). Perf is reported on the quality dataset.
comparison: the native model below a benchmark's floor or a Task without an Acc check), Red (Acc worse than
native beyond its margin), Yellow (an Acc difference not shown either way), Green (quality passes with available
benchmark timings). Historical fixed-workload reports retain their original performance lights.

### Performance

Expand All @@ -92,29 +92,33 @@ conversion parity only. The family's checks keep their documented coverage and t

Natural evaluation workloads (`both`) provide timings from the **same outputs used for quality**. Their
paired geometric speedup and total-time ratio are descriptive; the 90% interval is across dataset units,
with seeds clustered by problem, not a repeated-run stability interval. Every requested response must be
present, valid, paired, warmed, and at matching effective precision and declared task-call boundaries.
Actual work is compared per sample, so different samples may have different lengths. An unmatched
workload reports its natural-task time ratio and the reason equal-work acceleration is unavailable;
matched subsets never hide failures or shorter outputs. `max_tokens` alone is not work evidence. No
mandatory second, forced-length suite is added for variable-output families.
with seeds clustered by problem, not a repeated-run stability interval. The report shows each side's p50,
effective precision, successful timing coverage, and observed work differences.
Generation length and work comparability do not decide whether benchmark timings are measured.
Inputs excluded by Acc as exceeding bundle capacity are excluded from both timing sides too; their
count and the original attempted request counts remain visible. Other failed or missing requests
remain in coverage and produce a `partial` timing measurement when both sides have timings.
A missing workload or a side without any valid timing is an error. `max_tokens` alone is not work evidence.
No mandatory second, forced-length suite is added for variable-output families.

Models with quality benchmarks run **only their required quality workloads**. For example, Qwen uses
`mmlu-0shot`, and image/video models use their configured quality datasets. Catalog, near-capacity,
and informational replay checks are not extra default workloads. Each side answers each selected
problem once, with excluded warmup. The same profiling responses supply Acc, Native/TRTMC task-call
p50, and AIPerf client metrics. Multiple required benchmarks remain separate, labelled datasets.
Dataset timing completeness and work comparability are checked; timing results are measurements,
not repeated-run performance acceptance gates. Quality thresholds remain unchanged. Failed or
Dataset timing completeness and work comparability are recorded as observations; timing results
are measurements, not repeated-run performance acceptance gates. Quality thresholds remain unchanged. Failed or
unpaired responses remain visible in each side's timing coverage rather than disappearing into a
matched subset. Models explicitly lacking a quality benchmark retain one configured performance
workload and conversion-parity evidence. The explicit `order-check` diagnostic and historical
fixed-workload reports keep their original statistics. Timing and generation phases hold `gpu_lock`.

Configuration now uses a flat `performance` policy and opt-in `service_metrics`; reports use `performance`
and `service_metrics` under schema `trtmc.qualification/v2`. Earlier tiered configurations and reports
are normalized on read, including rejudge, without maintaining another execution path. Rejudge never
promotes descriptive dataset results to an acceptance gate. `torch.compile` remains an optional labelled
are normalized on read, including rejudge, without maintaining another execution path. Rejudge without an
environment refreshes benchmark timings from saved execution records while
preserving the recorded Acc entries and gates. It never promotes descriptive dataset results to
an acceptance gate. `torch.compile` remains an optional labelled
reference. Service metrics (client latency, throughput, load sweeps) are opt-in and do not affect the
verdict; the prototype's single execution lane and buffered SSE do not measure token TTFT/ITL.

Expand Down
124 changes: 124 additions & 0 deletions apps/aiperf_qual/tests/test_benchmark_perf.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0

import copy
import json
from types import SimpleNamespace

import pytest

from trtmc_aiperf_qual import benchmark_perf, execution, judge


def row(index, ms, tokens=1, text="C", **extra):
return {"metadata": {"session_num": index, "conversation_id": f"session_{index:06d}",
"benchmark_phase": "profiling"}, "status": 200,
"payload": {"prompt": str(index)}, "responses": [{"text": json.dumps({
"trtmc_timing": {"model_call_ms": ms},
"trtmc_observation": {"output_tokens": tokens, "text": text}})}], **extra}


def rejected(index):
return row(index, None, status=422, error={"message": json.dumps({"error": {
"code": "backend_rejected_request", "message": "prompt exceeds the prefill profile"}})})


def capture(out, candidate, native, accuracy=None):
evidence = execution.Session(out, {}, lambda: 0,
lambda obs: judge.work_signature("generate", obs),
lambda c, n: judge.work_check({"work": [c]}, {"work": [n]}) is None)
for side, records in (("candidate", candidate), ("reference", native)):
directory = out / side
directory.mkdir(parents=True)
(directory / "profile_export_raw.jsonl").write_text("".join(json.dumps(r) + "\n" for r in records))
evidence.record(SimpleNamespace(directory=directory, exit_code=0, raw_records=lambda: records), {
"name": "mmlu-0shot", "role": "both", "warmup": 1, "gpu_busy_percent": 0,
"expected_requests": len(records), "identity": {
"side": side, "precision": "fp16", "concurrency": 1, "timing_scope": "task-call-wall"}})
quality = accuracy or [{"suite": "mmlu-0shot", "source": "absolute", "status": "pass"}]
return {"model": "demo", "performance_source": "quality", "accuracy": quality,
"performance": evidence.natural_performance(quality),
"execution": {"records": str(out / "execution.jsonl")}}


def verdict(report):
return judge.verdict(report, expected_suites=["mmlu-0shot"], expected_modes=0)


def test_different_answer_lengths_are_descriptive_not_a_second_performance_gate(tmp_path):
report = capture(tmp_path, [row(0, 40, 4, "C. i")], [row(0, 20, 2)])
perf, = report["performance"]
assert verdict(report) == {"acc": "pass", "perf": "measured", "category": "measured", "lights": {}}
assert not perf["comparable"] and perf["matched_pairs"] == 0
assert perf["candidate"]["p50_ms"] == 40 and perf["reference"]["p50_ms"] == 20
assert not perf["gate"] and perf["measurement_status"] == "measured"


def test_capacity_exclusions_use_the_accuracy_scope_on_both_sides(tmp_path):
report = capture(tmp_path, [row(0, 40), rejected(1)], [row(0, 20), row(1, 200)], [
{"suite": "mmlu-0shot", "source": "absolute", "status": "pass", "out_of_capacity": 1}])
perf, = report["performance"]
assert verdict(report)["perf"] == "measured"
assert perf["complete"] and perf["out_of_capacity"] == 1
assert perf["reference"]["p50_ms"] == 20
assert all(perf[s]["requests"] == perf[s]["valid_requests"] == 1 for s in ("candidate", "reference"))
assert all(perf[s]["attempted_requests"] == 2 for s in ("candidate", "reference"))


def test_partial_failed_workload_reports_coverage_and_available_timings(tmp_path):
report = capture(tmp_path, [row(0, 40), row(1, None, status=500)], [row(0, 20), row(1, 200)])
perf, = report["performance"]
assert verdict(report)["perf"] == "partial" and verdict(report)["category"] == "measured"
assert not perf["complete"] and perf["candidate"]["valid_requests"] == 1
assert perf["reference"]["p50_ms"] == 110
assert perf["measurement_status"] == "partial"


def test_no_candidate_timing_remains_an_error(tmp_path):
report = capture(tmp_path, [row(0, None)], [row(0, 20)])
assert verdict(report)["perf"] == "error"
assert report["performance"][0]["measurement_status"] == "unavailable"


def test_refresh_old_exports_preserves_accuracy_and_raw_responses(tmp_path):
report = capture(tmp_path, [row(0, 40), rejected(1)], [row(0, 20), row(1, 200)], [
{"suite": "mmlu-0shot", "source": "absolute", "status": "pass", "out_of_capacity": 1,
"gate": {"margin": 1.0}, "metrics": {"trtmc_score": 80, "native_score": 80}}])
path = tmp_path / "execution.jsonl"
old = [json.loads(line) for line in path.read_text().splitlines()]
for batch in old:
for r in batch["records"]:
r.pop("capacity_rejection")
path.write_text("".join(json.dumps(batch) + "\n" for batch in old))
before = {p: p.read_bytes() for p in tmp_path.rglob("*.jsonl")}
quality = copy.deepcopy(report["accuracy"])
updated = benchmark_perf.refresh(tmp_path, report)
assert updated["accuracy"] == quality
assert updated["performance"][0]["reference"]["p50_ms"] == 20
assert verdict(updated)["perf"] == "measured"
assert all(p.read_bytes() == data for p, data in before.items())
assert benchmark_perf.refresh(tmp_path, updated) == updated


def test_capacity_count_mismatch_is_not_silently_filtered(tmp_path):
with pytest.raises(ValueError, match="capacity rejections"):
capture(tmp_path, [row(0, 40), rejected(1)], [row(0, 20), row(1, 200)], [
{"suite": "mmlu-0shot", "status": "pass", "out_of_capacity": 2}])


def test_capacity_excluded_problem_leaves_all_seed_repetitions(tmp_path):
capture(tmp_path, [row(0, 40), rejected(1)], [row(0, 20), row(1, 200)], [
{"suite": "mmlu-0shot", "status": "pass", "out_of_capacity": 1}])
batches = [json.loads(line) for line in (tmp_path / "execution.jsonl").read_text().splitlines()]
more = copy.deepcopy(batches)
for batch in more:
for r in batch["records"]:
r["request_sha"] += "-seed-two"
r["capacity_rejection"] = None
r["valid"] = True
r["model_call_ms"] = 80
perf = execution.paired_dataset("mmlu-0shot", [batches[0], more[0]], [batches[1], more[1]],
lambda c, n: True, capacity_exclusions=1)
assert perf["complete"] and perf["pairs"] == 2 and perf["out_of_capacity"] == 1
assert all(perf[s]["requests"] == 2 and perf[s]["attempted_requests"] == 4
for s in ("candidate", "reference"))
3 changes: 2 additions & 1 deletion apps/aiperf_qual/tests/test_execution.py
Original file line number Diff line number Diff line change
Expand Up @@ -268,7 +268,8 @@ def test_quality_measurement_verdict_does_not_claim_repeated_performance_gate(tm
result = judge.verdict(base, expected_suites=["evaluation"], expected_modes=0)
assert result == {"acc": "pass", "perf": "measured", "lights": {}, "category": "measured"}
item["complete"] = False
assert judge.verdict(base, expected_suites=["evaluation"], expected_modes=0)["category"] == "error"
partial = judge.verdict(base, expected_suites=["evaluation"], expected_modes=0)
assert partial["perf"] == "partial" and partial["category"] == "measured"
item["complete"] = True
base["accuracy"].append({"suite": "another-required-dataset", "source": "absolute", "status": "pass"})
assert judge.verdict(base, expected_suites=["evaluation"], expected_modes=0)["perf"] == "error"
Expand Down
7 changes: 5 additions & 2 deletions apps/aiperf_qual/trtmc_aiperf_qual/accuracy_recovery.py
Original file line number Diff line number Diff line change
Expand Up @@ -130,6 +130,8 @@ def aligned_batches(recorded: list[dict]) -> list[dict]:
if unit is None:
raise ValueError("original evaluation unit is missing from the saved execution batch")
rows.append({**row, "sample_id": identity, "unit_id": unit})
if batch["identity"].get("side") == "candidate":
rows[-1]["capacity_rejection"] = absolute.capacity_rejection(source)
aligned.append({**batch, "records": rows, "alignment": ALIGNMENT})
return aligned

Expand Down Expand Up @@ -181,13 +183,14 @@ def recover(out: Path, model: dict, report: dict, archive: SelectionArchive) ->
{"work": [mine] if mine is not None else []}, {"work": [theirs] if theirs is not None else []}) is None
session = execution.Session(out, {}, lambda: None, lambda value: value, same_work, batches=aligned)
performance = [item for item in report.get("performance", []) if item.get("kind") != "natural_dataset"]
performance.extend(session.natural_performance())
accuracy = [updates.get(entry["suite"], entry) for entry in report.get("accuracy", [])]
performance.extend(session.natural_performance(accuracy))
# Validate the whole report before publishing any corrected evidence.
for path, grades in pending_exports:
partial = path.with_suffix(".tmp")
partial.write_text("".join(json.dumps(row, ensure_ascii=False) + "\n" for row in grades))
partial.replace(path)
result = {**report, "accuracy": [updates.get(entry["suite"], entry) for entry in report.get("accuracy", [])],
result = {**report, "accuracy": accuracy,
"performance": performance}
result["accuracy_alignment"] = {"version": ALIGNMENT, "suites": sorted(updates), "runs": evidence,
"original_responses_reused": True}
Expand Down
62 changes: 62 additions & 0 deletions apps/aiperf_qual/trtmc_aiperf_qual/benchmark_perf.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
"""Refresh benchmark timing reports from saved responses, without inference or regrading."""

from __future__ import annotations

import hashlib
import json
from pathlib import Path

from . import absolute, execution, judge
from .aiperf_runner import AiperfRun


def refresh(out: Path, report: dict) -> dict:
if report.get("performance_source") != "quality":
return report
recorded_path = Path((report.get("execution") or {}).get("records", "execution.jsonl"))
path = out / recorded_path.name
if not path.is_file():
if not report.get("execution") and not any(item.get("out_of_capacity") for item in report.get("accuracy", [])):
# Older aggregate-only reports can reapply the measurement verdict,
# but cannot reconstruct a different sample selection.
return report
raise ValueError(f"{out}: recorded benchmark execution is unavailable")
batches, superseded = [], set()
for line in path.read_text().split("\n"):
if not line.strip():
continue
batch = json.loads(line)
if batch.get("event") == "supersede_failed_attempt":
superseded.update(batch["batch_ids"])
elif "records" in batch:
batches.append(batch)
batches = [batch for batch in batches
if not batch.get("superseded") and batch["batch_id"] not in superseded]
excluded = {item["suite"] for item in report.get("accuracy", [])
if item.get("out_of_capacity") and item.get("status") != "error"}
# Older execution exports omit rejection reasons. Read their original raw
# error records only when accuracy explicitly excluded capacity rejections.
raw = {}
for batch in batches:
if batch["workload"] not in excluded or batch["identity"].get("side") != "candidate":
continue
for row in batch["records"]:
if row["output_valid"] or "capacity_rejection" in row:
continue
ref = row["output_ref"]
directory = ref["aiperf_run"]
if directory not in raw:
raw[directory] = AiperfRun(Path(directory), batch["aiperf_exit"], []).raw_records()
source = raw[directory][ref["record_index"]]
payload = source.get("payload") or {}
sha = hashlib.sha256(json.dumps(payload, sort_keys=True, separators=(",", ":")).encode()).hexdigest()
if row["request_sha"] != sha or str(row["sample_id"]) != source["metadata"].get("conversation_id"):
raise ValueError("saved capacity rejection does not match its execution identity")
row["capacity_rejection"] = absolute.capacity_rejection(source)
same_work = lambda mine, theirs: judge.work_check( # noqa: E731
{"work": [mine] if mine is not None else []}, {"work": [theirs] if theirs is not None else []}) is None
evidence = execution.Session(out, {}, lambda: None, lambda value: value, same_work, batches=batches)
performance = [item for item in report.get("performance", []) if item.get("kind") != "natural_dataset"]
return {**report, "performance": [*performance, *evidence.natural_performance(report.get("accuracy", []))]}
11 changes: 11 additions & 0 deletions apps/aiperf_qual/trtmc_aiperf_qual/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -183,6 +183,17 @@ def rejudge_reports(outs: Sequence[Path], environment=None, *, selection_cache:
continue
result = compat.report(json.loads(path.read_text()))
model = compat.configuration(json.loads((out / "model.json").read_text()))
if selection_cache is None and environment is None and result.get("performance_source") == "quality":
from .benchmark_perf import refresh

result = refresh(out, result)
result["verdict"] = judge.verdict(result, expected_suites=list(expected_suites(model)), expected_modes=0)
preserve_original(out)
result["rejudged"] = {"time": time.time(), "original": ORIGINAL_REPORT,
"benchmark_timings_refreshed": True, "accuracy_contract_preserved": True}
write_report(out, result)
print(json.dumps({"out": str(out), **result["verdict"]}))
continue
if selection_cache is not None:
preserve_original(out)
result = recover(out, model, result, archive)
Expand Down
Loading
Loading