Skip to content

fix: compare only complete browser benchmark costs - #173

Open
rudycelekli wants to merge 1 commit into
Merit-Systems:mainfrom
rudycelekli:fix/compare-complete-benchmark-costs
Open

rudycelekli wants to merge 1 commit into
Merit-Systems:mainfrom
rudycelekli:fix/compare-complete-benchmark-costs

Conversation

@rudycelekli

Copy link
Copy Markdown

Summary

Do not report cost improvements when a browser benchmark has incomplete cost measurements.

The supported benchmark producer sums known step.completed costs and sets costComplete: false if any step omits usage.costUsd. Its report and dashboard already show those partial totals with ~. The comparison CLI and dashboard improvement helper discarded that flag, so a partial $1 measurement versus a complete $2 baseline appeared as a precise 50% saving.

Preserve the approximation marker in the comparison CLI's task and total costs. Compute task deltas and comparable aggregate spend only from shared passing tasks with complete costs in both variants. The dashboard excludes those incomplete pairs from cost improvements and their mean. Duration comparisons continue to use shared passing tasks.

Native reproduction and controls

The regression invokes the real measureWorkerTask producer with supported Eve step events, including a completed step without optional cost usage. It feeds the resulting metrics through the public benchmark JSON schema and runs the actual scripts/compare-browser-benchmarks.ts CLI as a bounded child process. It also calls the actual helper used by the dashboard's task and mean comparisons.

  • Unchanged source: 2 failures / 3 existing controls pass. The helper reports -0.5 for an incomplete cost, and the real CLI displays a precise $1.000000 and (-50.0%).
  • After: 11/11 focused comparison and worker-event tests pass. Complete costs still report the true 50% saving; incomplete costs remain marked, produce no cost delta and keep their duration comparison. Existing complete, averaged and unsuccessful-pair controls pass.
  • pnpm check --concurrency=1: 95 files / 878 tests, all six lint/types/tests/format/Knip tasks pass.
  • pnpm build: passes using owned placeholder database/Kernel settings.
  • Standalone dashboard: next typegen evals/browser/dashboard and tsc --noEmit -p evals/browser/dashboard/tsconfig.json pass. Root types exclude this app, so its caller compatibility was checked directly.

Local runtime validation was blocked by host Docker cleanup: an earlier unchanged runtime runner and copied production workflow tools displayed two mock-model evals and 13 gates passing, then exited 1 at the shipped 180-second cleanup timeout. Exact owned docker ps cleanup also independently timed out. This branch did not repeat that blocked local run.

The unchanged declared Checks workflow completed successfully on the exact signed source head e303ad513b21e3b9468172f73499717e73c2a2ab. Hosted pnpm check passed 95 files / 878 tests and all six check tasks. Hosted pnpm test:runtime completed successfully: both isolated mock-model workflow evals passed all 13 gates, including cleanup. The run used Node 24.21.0 from the declared Node 24 version and pnpm 11.24.0 with the frozen lockfile. The PR merge checkout tree matches the reviewed source head; no workflow, dependency or runtime fixture changes were made.

Scope and limitations

One shared cost-completeness cause across the CLI and dashboard, with a short contract clarification in the benchmark README. No pricing, token measurement, model, benchmark schema, workflow, dependency or runtime changes.

These tests execute supported event measurement, real JSON validation, the real comparison CLI and the dashboard helper. They do not run models, charge a provider, run a live browser benchmark or claim a mounted dashboard test. Current open PRs, issues and recent merged/closed PR bodies were screened for overlap; none address this cost-completeness cause.

AI assistance was used for investigation, implementation, tests and review under the submitting account. Signed DCO commit included.

Signed-off-by: Rudy Celekli <rudy@gradiahq.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant