Repository navigation
fix: compare only complete browser benchmark costs - #173
Open
rudycelekli wants to merge 1 commit into
Open
rudycelekli wants to merge 1 commit into
rudycelekli wants to merge 1 commit into
Conversation
Signed-off-by: Rudy Celekli <rudy@gradiahq.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Do not report cost improvements when a browser benchmark has incomplete cost measurements.
The supported benchmark producer sums known
step.completedcosts and setscostComplete: falseif any step omitsusage.costUsd. Its report and dashboard already show those partial totals with~. The comparison CLI and dashboard improvement helper discarded that flag, so a partial $1 measurement versus a complete $2 baseline appeared as a precise 50% saving.Preserve the approximation marker in the comparison CLI's task and total costs. Compute task deltas and comparable aggregate spend only from shared passing tasks with complete costs in both variants. The dashboard excludes those incomplete pairs from cost improvements and their mean. Duration comparisons continue to use shared passing tasks.
Native reproduction and controls
The regression invokes the real
measureWorkerTaskproducer with supported Eve step events, including a completed step without optional cost usage. It feeds the resulting metrics through the public benchmark JSON schema and runs the actualscripts/compare-browser-benchmarks.tsCLI as a bounded child process. It also calls the actual helper used by the dashboard's task and mean comparisons.-0.5for an incomplete cost, and the real CLI displays a precise$1.000000and(-50.0%).pnpm check --concurrency=1: 95 files / 878 tests, all six lint/types/tests/format/Knip tasks pass.pnpm build: passes using owned placeholder database/Kernel settings.next typegen evals/browser/dashboardandtsc --noEmit -p evals/browser/dashboard/tsconfig.jsonpass. Root types exclude this app, so its caller compatibility was checked directly.Local runtime validation was blocked by host Docker cleanup: an earlier unchanged runtime runner and copied production workflow tools displayed two mock-model evals and 13 gates passing, then exited 1 at the shipped 180-second cleanup timeout. Exact owned
docker pscleanup also independently timed out. This branch did not repeat that blocked local run.The unchanged declared Checks workflow completed successfully on the exact signed source head
e303ad513b21e3b9468172f73499717e73c2a2ab. Hostedpnpm checkpassed 95 files / 878 tests and all six check tasks. Hostedpnpm test:runtimecompleted successfully: both isolated mock-model workflow evals passed all 13 gates, including cleanup. The run used Node 24.21.0 from the declared Node 24 version and pnpm 11.24.0 with the frozen lockfile. The PR merge checkout tree matches the reviewed source head; no workflow, dependency or runtime fixture changes were made.Scope and limitations
One shared cost-completeness cause across the CLI and dashboard, with a short contract clarification in the benchmark README. No pricing, token measurement, model, benchmark schema, workflow, dependency or runtime changes.
These tests execute supported event measurement, real JSON validation, the real comparison CLI and the dashboard helper. They do not run models, charge a provider, run a live browser benchmark or claim a mounted dashboard test. Current open PRs, issues and recent merged/closed PR bodies were screened for overlap; none address this cost-completeness cause.
AI assistance was used for investigation, implementation, tests and review under the submitting account. Signed DCO commit included.