Skip to content

docs(evals): bind current Flow release performance#270

Merged
abrichr merged 1 commit into
mainfrom
eval/current-flow-local-20260718
Jul 19, 2026
Merged

docs(evals): bind current Flow release performance#270
abrichr merged 1 commit into
mainfrom
eval/current-flow-local-20260718

Conversation

@abrichr

@abrichr abrichr commented Jul 19, 2026

Copy link
Copy Markdown
Member

Outcome

Adds a reproducible, local-only performance readout bound to the exact published openadapt-flow v1.16.1 wheel and release tag.

  • 3 trials per arm per condition across clean, theme, and label-rename surfaces
  • compiled replay versus steelmanned positional and name-scoped Playwright controls
  • one shared screenshot/OCR effect oracle for exact note, Triage row, intended patient, and wrong-target writes
  • steady action-loop and fresh-browser end-to-end median/p95
  • three record/compile setup trials
  • explicit correct / silent incorrect / wrong action / over-halt / halt-error taxonomy
  • zero model calls, zero model cost, no VM/provider mutations

Exact binding

  • Flow tag: v1.16.1
  • Flow release commit: 113ce992b491576d77236f495b983165ce7a63bd
  • Published wheel SHA-256: c7073283475e7ae722db2478b499d962364851104e09edb71f254ba21c1310cd
  • Evals base: 7629ba5a2447c919e499e88e9b4eedbc80c8f3ab
  • Runner SHA-256: ac58c0b9a02cfc991144a1f5c7a2814c6b6b748bcba27b87f8e1027ee575970f

Result

Compiled effect oracle: clean 3/3, theme 3/3, rename 3/3. Theme is deliberately reported as 3/3 over-halt because the intended effect was present while region_stable left Replayer unsuccessful. Both selector controls pass clean/theme 3/3 and fail label rename loudly 0/3 before mutation. All cells have zero silent incorrect success, wrong action, model calls, and model cost.

This is deterministic runtime overhead/robustness evidence, not a zero-shot comparison. The report explicitly leaves current Flow versus zero-shot unmeasured and records the missing WAA evaluator/hybrid wiring needed before any paid run would be valid.

Verification

  • ruff check and ruff format --check on runner/tests
  • 9 focused tests passed (test_current_flow_local_benchmark.py + existing performance-report tests)
  • final artifact/source assertions passed
  • no screenshots, copied app source, AGPL benchmark material, or package artifact changes

@abrichr
abrichr marked this pull request as ready for review July 19, 2026 02:47
@abrichr
abrichr merged commit 8b00d3c into main Jul 19, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant