Skip to content

feat(experiments): controlled baseline-vs-candidate eval harness for … - #307

Open
Matthew (theworker02) wants to merge 1 commit into
microsoft:mainfrom
theworker02:devin/skillopt-eval-harness
Open

Matthew (theworker02) wants to merge 1 commit into
microsoft:mainfrom
theworker02:devin/skillopt-eval-harness

Conversation

@theworker02

Copy link
Copy Markdown

…#132

Adds experiments/skillopt-eval: a pluggable-runner experiment that splits scenarios into seeded train/validation/test sets, gates bounded skill edits behind strict held-out validation improvement (fail-closed), and emits paired McNemar + bootstrap stats with full provenance as raw_results.csv / results.json / RESULTS.md plus raw trajectories.

Runners: 'mock' (deterministic offline dry-run, no keys/network) and 'superpowers' (wraps the existing isolated-overlay adapter at a pinned SHA; baseline and candidate share harness/model/settings by construction). Deterministic offline tests cover gate accept/reject, fail-closed errors, learning-rate bounding, source non-mutation, artifacts, and A/A no-claim.

…icrosoft#132

Adds experiments/skillopt-eval: a pluggable-runner experiment that
splits scenarios into seeded train/validation/test sets, gates bounded
skill edits behind strict held-out validation improvement (fail-closed),
and emits paired McNemar + bootstrap stats with full provenance as
raw_results.csv / results.json / RESULTS.md plus raw trajectories.

Runners: 'mock' (deterministic offline dry-run, no keys/network) and
'superpowers' (wraps the existing isolated-overlay adapter at a pinned
SHA; baseline and candidate share harness/model/settings by construction).
Deterministic offline tests cover gate accept/reject, fail-closed errors,
learning-rate bounding, source non-mutation, artifacts, and A/A no-claim.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@Yif-Yang

Copy link
Copy Markdown
Contributor

Reviewed head ce4e6f7c5629e4a15bb4032ab0f7064151a2fd81 against current main. This provides useful candidate-selection orchestration around the existing adapter and evalkit, but I am holding merge for these production-path gaps:

  • The live runner gives baseline, candidate and successive attempts the same scenario/seed/pin trajectory filename. Candidate output overwrites baseline evidence, and the stored condition/split are empty. Baseline validation records are also absent from raw_results.csv. Please preserve a separately labeled artifact for every invocation and retain the reference used for each gate decision.
  • With optimize_steps=0, two runs of the same unchanged baseline can produce optimized=false yet quality_improved=true. An offline provider-boundary reproduction with one discordant pair returned exact McNemar p=1 but still emitted the candidate-improvement verdict. The report needs to preserve the declared A/A no-claim invariant and distinguish observed run variation from a candidate effect.
  • A validation batch with one pass and one timeout is accepted over a zero baseline. The timeout is counted as a failed row, but this differs from the candidate-level guarantee in test_validation_fails_closed_on_runner_error. Please align the implementation, coverage and documented fail-closed contract.

The existing focused tests pass: 150 tests plus 84 subtests. The independent reproductions exercised the actual Superpowers adapter, local pinned synthetic checkout, judges and experiment loop; only the external Claude process was substituted. No paid model calls or model-efficiency claims are involved. These require author revisions, so this is not a merge recommendation.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants