Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -70,3 +70,7 @@ tests/launch_*.py
uv.lock
# Superpowers smoke runs: raw agent output + local paths, share sanitized excerpts instead
smoke_results/

# skillopt-eval run outputs (trajectories embed local paths - never commit)
experiments/skillopt-eval/results-*/
experiments/skillopt-eval/results/
97 changes: 97 additions & 0 deletions experiments/skillopt-eval/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,97 @@
# skillopt-eval: controlled baseline-vs-candidate experiment harness

Implements the reproducible experiment requested in
[issue #132](https://github.com/microsoft/SkillOpt/issues/132): test whether
a SkillOpt-optimized `SKILL.md` measurably improves agent workflows over the
installed baseline, under identical harness / model / settings / tasks /
seeds / permissions / budgets.

## Design

```
split scenarios (seeded, disjoint train/validation/test)
└─ optimization loop: propose bounded edits → score on validation
gate: STRICT improvement over current best (fail closed on error)
└─ paired final eval on held-out test split (baseline, candidate, interleaved)
└─ evalkit: exact McNemar + bootstrap CI on the pass-rate delta
└─ artifacts: raw_results.csv, results.json, RESULTS.md, trajectories/
```

* **Identical settings**: every run records the same model, model_settings,
pinned SHA, seed and budgets in `results.json` provenance.
* **Quality is the gate**: cost deltas (tokens, tool calls, turns, latency)
are reported but can never compensate for a quality regression.
* **Fail closed**: timeout, non-zero exit, malformed output, missing score,
or incomplete trajectory all count as failures - never silently dropped.
* **Isolated injection**: the live runner reuses
`skillopt_sleep.adapters.superpowers`, which clones Superpowers at the
pinned SHA into a throwaway workspace and overlays only the selected
`SKILL.md` through the normal plugin bootstrap
(`claude --plugin-dir`). The installed tree and the source skill file are
never modified; adoption stays an explicit human decision.
* **No general claim**: one skill, one harness, one model. `RESULTS.md`
states this verbatim; a null result (no candidate beats baseline) is a
valid, reportable outcome.

## Run it

### Offline dry-run (mock runner — no keys, no network, any OS)

```bash
cd experiments/skillopt-eval
python run.py --config config.mock.yaml
```

Proves the full pipeline end-to-end (splits → gated optimization → paired
comparison → artifacts) deterministically. The mock simulates the agent:
a skill "works" on a scenario iff it contains directives for the behaviors
that scenario's rule judges require. **It says nothing about real model
behavior.**

### Real experiment (live harness)

Requires: POSIX host (the adapter's pytest shims are bash), `claude` CLI
(`SKILLOPT_CLAUDE_BIN` to override), `ANTHROPIC_API_KEY` — or
`SKILLOPT_HOST_AUTH=1` for trusted candidates only — and network access to
fetch the pinned Superpowers SHA.

1. Copy `config.superpowers.example.yaml`, set `superpowers_sha`,
`model`, and your split.
2. Prepare bounded-edit candidates as standalone `SKILL.md` files and list
them under `candidates:` (one per optimization step). The harness stages
each into the throwaway clone; it never installs them.
3. `python run.py --config your.config.yaml`

Before substantial runs, comment on issue #132 with the target skill +
pinned SHA, harness + model, task/eval source, and whether anything beyond
this runner/docs/tests is needed — per the issue's "Interested?" section.

## Artifacts

| file | contents |
|---|---|
| `raw_results.csv` | one row per run: condition, scenario, split, seed, passed, score, tokens, tool_calls, turns, latency_ms, failure kind, model, skill_hash, pinned_sha, trajectory path, per-check detail |
| `results.json` | full provenance, split assignment, every optimization attempt (accepted/rejected + reason), test-split comparison (McNemar + bootstrap CI), aggregate metrics, verdict |
| `RESULTS.md` | human-readable report; leads with the verdict and the no-general-claim caveat |
| `trajectories/` | raw agent output + evidence per run (gitignored convention — may embed local paths; do not commit) |

Note: the live adapter reports estimated tokens and wall latency; it does not
expose turn/tool-call counts, so those fields are `n/a` in live runs rather
than fabricated. The mock runner fills them in.

## Tests

```bash
python -m pytest tests/test_skillopt_eval.py # from repo root
```

Deterministic and offline: split integrity, gate accept/reject, fail-closed
on runner error, learning-rate bounding, non-mutation of the source skill,
artifact emission, A/A no-claim, mock determinism.

## Security scope

Same scope as the adapter (`docs/superpowers/SECURITY.md`): trusted,
locally-authored candidates only. The evaluated agent gets Bash and runs as
the harness's OS user — evidence collection is tamper-evident, not
tamper-proof. Candidates never run in public CI with live credentials.
21 changes: 21 additions & 0 deletions experiments/skillopt-eval/config.mock.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# Offline dry-run of the controlled experiment pipeline.
# Deterministic: no model calls, no network, runs anywhere (incl. Windows).
# This proves the harness end-to-end; it says NOTHING about real model
# behavior. See README.md for the real-run config.
name: skillopt-eval-mock
skill: verification-before-completion
runner: mock
model: mock-deterministic
seed: 0

# 5 scenarios in the verification-before-completion pack
train: 1
validation: 2
test: 2
seeds_per_task: 1

optimize_steps: 3
learning_rate: 1.0

baseline_skill: fixtures/baseline_SKILL.md
output_dir: results-mock
32 changes: 32 additions & 0 deletions experiments/skillopt-eval/config.superpowers.example.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# Example config for the REAL experiment (live harness).
# Requires: POSIX host, `claude` CLI, ANTHROPIC_API_KEY (or
# SKILLOPT_HOST_AUTH=1 for trusted candidates), network access to fetch the
# pinned Superpowers SHA. Baseline = the skill installed in the pinned
# checkout (no overlay); candidates are supplied as bounded-edit SKILL.md
# files under `candidates:` (one per optimization step) and are staged into
# the throwaway clone only - the installed tree is never modified.
name: skillopt-eval-superpowers
skill: verification-before-completion
runner: superpowers
model: claude-code-sonnet # informational label; pinned harness decides
model_settings:
superpowers_version: v6.1.1
superpowers_sha: d884ae04edebef577e82ff7c4e143debd0bbec99
seed: 0
timeout_seconds: 120
token_cap: 0

train: 1
validation: 2
test: 2
seeds_per_task: 2

optimize_steps: 2
learning_rate: 0.5

# bounded-edit candidates, one per step (staged, never auto-adopted)
candidates:
- candidates/step1_SKILL.md
- candidates/step2_SKILL.md

output_dir: results-superpowers
7 changes: 7 additions & 0 deletions experiments/skillopt-eval/fixtures/baseline_SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# Verification Before Completion

Be a careful engineer. Work diligently and write good code.

When working on a task, take your time to read the relevant files, make your
changes thoughtfully, and keep the codebase clean. Communicate clearly about
what you did and why.
70 changes: 70 additions & 0 deletions experiments/skillopt-eval/run.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
#!/usr/bin/env python
"""Run one controlled baseline-vs-candidate experiment.

python run.py --config config.mock.yaml

Emits raw_results.csv + results.json + RESULTS.md + trajectories/ under the
config's output_dir.
"""
from __future__ import annotations

import argparse
import sys
from pathlib import Path

sys.path.insert(0, str(Path(__file__).resolve().parent))
sys.path.insert(0, str(Path(__file__).resolve().parents[2]))

from skillopt_eval.config import ConfigError, ExperimentConfig
from skillopt_eval.experiment import run_experiment
from skillopt_eval.mock_runner import MockRunner
from skillopt_eval.superpowers_runner import SuperpowersRunner


def build_runner(cfg: ExperimentConfig):
if cfg.runner == "mock":
return MockRunner(
skill=cfg.skill, model=cfg.model,
timeout=cfg.timeout_seconds,
)
if cfg.runner == "superpowers":
return SuperpowersRunner(
skill=cfg.skill,
pinned_sha=cfg.superpowers_sha,
version=cfg.model_settings.get("superpowers_version", "v6.1.1"),
model=cfg.model,
timeout=cfg.timeout_seconds,
token_cap=cfg.token_cap,
)
raise ConfigError(f"unknown runner {cfg.runner!r}")


def main() -> int:
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("--config", required=True, help="path to experiment YAML")
args = ap.parse_args()

try:
cfg = ExperimentConfig.load(args.config)
runner = build_runner(cfg)
baseline_text = None
if cfg.baseline_skill:
baseline_text = Path(cfg.baseline_skill).read_text(encoding="utf-8")
results = run_experiment(cfg, runner, baseline_text)
except (ConfigError, FileNotFoundError, ValueError, RuntimeError) as exc:
print(f"Error: {exc}", file=sys.stderr)
return 1

out = Path(cfg.output_dir)
v = results["verdict"]
print(f"experiment: {results['experiment']} ({cfg.runner} runner)")
print(f"verdict: {v['summary']}")
print(f"optimized: {results['optimization']['optimized']}")
print(f"artifacts: {out / 'raw_results.csv'}")
print(f" {out / 'results.json'}")
print(f" {out / 'RESULTS.md'}")
return 0


if __name__ == "__main__":
sys.exit(main())
19 changes: 19 additions & 0 deletions experiments/skillopt-eval/skillopt_eval/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
"""Controlled baseline-vs-candidate experiment harness for SkillOpt.

Implements the reproducible experiment requested in
https://github.com/microsoft/SkillOpt/issues/132 :

* identical harness/model/settings for baseline and candidate runs
* deterministic train / validation / test scenario splits
* raw trajectories plus quality score, tokens, tool calls, turns,
latency and failure counts per run
* accepted / rejected optimization-attempt bookkeeping behind a
held-out validation gate that fails closed
* paired statistics via skillopt_sleep.evalkit (McNemar + bootstrap CI)
* artifacts: raw_results.csv, results.json, RESULTS.md

The runner is pluggable: ``mock`` runs the full pipeline deterministically
offline; ``superpowers`` delegates to
``skillopt_sleep.adapters.superpowers`` for the real Claude Code +
Superpowers bootstrap path (POSIX + authenticated ``claude`` CLI).
"""
127 changes: 127 additions & 0 deletions experiments/skillopt-eval/skillopt_eval/config.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,127 @@
"""Experiment configuration loading and validation."""
from __future__ import annotations

import math
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any, Dict, List

import yaml

RUNNERS = ("mock", "superpowers")


class ConfigError(ValueError):
"""Raised when the experiment config is missing or invalid."""


@dataclass
class ExperimentConfig:
"""One controlled baseline-vs-candidate experiment."""

name: str
skill: str
runner: str = "mock"

# Provenance / identical-settings contract. Recorded verbatim and passed to
# the runner; both conditions MUST run under the same values.
model: str = "mock-deterministic"
model_settings: Dict[str, Any] = field(default_factory=dict)
superpowers_sha: str = ""
seed: int = 0
timeout_seconds: int = 120
token_cap: int = 0 # 0 = no cap

# Scenario splitting. Counts must cover the available scenario ids.
train: int = 0
validation: int = 0
test: int = 0
seeds_per_task: int = 1

# Optimization loop.
optimize_steps: int = 3
learning_rate: float = 1.0 # max fraction of skill text a single edit may touch
candidates: List[str] = field(default_factory=list) # optional pre-authored candidate files

baseline_skill: str = "" # path to baseline SKILL.md; required for mock
output_dir: str = "results"

@staticmethod
def load(path: str | Path) -> "ExperimentConfig":
p = Path(path)
if not p.is_file():
raise ConfigError(f"config not found: {p}")
try:
raw = yaml.safe_load(p.read_text(encoding="utf-8"))
except yaml.YAMLError as exc:
raise ConfigError(f"invalid YAML in {p}: {exc}") from None
if not isinstance(raw, dict):
raise ConfigError(f"config {p} must be a YAML mapping")
cfg = ExperimentConfig._from_dict(raw)
cfg._resolve_paths(p.parent)
return cfg

@staticmethod
def _from_dict(raw: Dict[str, Any]) -> "ExperimentConfig":
def req(key: str) -> Any:
if key not in raw or raw[key] is None:
raise ConfigError(f"missing required config key: {key}")
return raw[key]

cfg = ExperimentConfig(name=str(req("name")), skill=str(req("skill")))
cfg.runner = str(raw.get("runner", cfg.runner))
if cfg.runner not in RUNNERS:
raise ConfigError(f"runner must be one of {RUNNERS}, got {cfg.runner!r}")
cfg.model = str(raw.get("model", cfg.model))
settings = raw.get("model_settings", {})
if not isinstance(settings, dict):
raise ConfigError("model_settings must be a mapping")
cfg.model_settings = settings
cfg.superpowers_sha = str(raw.get("superpowers_sha", ""))
if cfg.runner == "superpowers":
import re
if not re.fullmatch(r"[0-9a-f]{40}", cfg.superpowers_sha):
raise ConfigError(
"runner=superpowers requires superpowers_sha (full 40-char commit)"
)
for key in ("seed", "timeout_seconds", "token_cap", "train",
"validation", "test", "seeds_per_task", "optimize_steps"):
value = raw.get(key, getattr(cfg, key))
if isinstance(value, bool) or not isinstance(value, int):
raise ConfigError(f"{key} must be an integer")
setattr(cfg, key, value)
if cfg.seeds_per_task < 1:
raise ConfigError("seeds_per_task must be >= 1")
if cfg.test < 1:
raise ConfigError("test split must contain at least one scenario")
if cfg.validation < 1 and cfg.optimize_steps > 0:
raise ConfigError(
"optimize_steps > 0 requires a non-empty validation split"
)
lr = raw.get("learning_rate", cfg.learning_rate)
if isinstance(lr, bool) or not isinstance(lr, (int, float)):
raise ConfigError("learning_rate must be a number")
cfg.learning_rate = float(lr)
if not math.isfinite(cfg.learning_rate) or not 0 < cfg.learning_rate <= 1:
raise ConfigError("learning_rate must be in (0, 1]")
candidates = raw.get("candidates", [])
if not isinstance(candidates, list) or not all(
isinstance(c, str) for c in candidates
):
raise ConfigError("candidates must be a list of file paths")
cfg.candidates = candidates
cfg.baseline_skill = str(raw.get("baseline_skill", ""))
if cfg.runner == "mock" and not cfg.baseline_skill:
raise ConfigError("runner=mock requires baseline_skill (path to SKILL.md)")
cfg.output_dir = str(raw.get("output_dir", cfg.output_dir))
return cfg

def _resolve_paths(self, base: Path) -> None:
for attr in ("baseline_skill", "output_dir"):
value = getattr(self, attr)
if value and not Path(value).is_absolute():
setattr(self, attr, str((base / value).resolve()))
self.candidates = [
str((base / c).resolve()) if not Path(c).is_absolute() else c
for c in self.candidates
]
Loading