Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,17 @@ All notable changes to the claude-plugins project will be documented in this fil

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). Entries are listed newest-first; each plugin section is treated as released when merged to `main`.

### code v1.15.1

#### Added
- `skills/guided-manual-qa/agents/openai.yaml` carries Codex display metadata, so Codex and Claude Code load the same `guided-manual-qa` skill directory.
- `guided-manual-qa` keeps a QA session alive across tool calls or worker turns: services run under a repository-supported or OS-supported owner that survives that boundary, the record says how to inspect and stop it, and a resumed session rereads the QA record and rechecks the head, owned processes, listeners, data target, and route instead of trusting earlier PIDs or ready checks.
- The QA record template gains a line for the interactive window or app owner, its settled route and control, and the last live verification time, plus a note under the services table to record each process owner and recheck the rows after a resume.

#### Changed
- Before a UI checkpoint, the agent opens the requested window or app itself, verifies the settled origin and a visible control owned by the route, and keeps it available; an unready window keeps the checkpoint pending as a setup limitation. After presenting a checkpoint the agent stops making tool calls until the human responds or asks for setup help.
- `SKILL.md` names the bundled launcher as `scripts/dist/launch-interactive-browser.mjs` relative to the skill directory and has the agent resolve its absolute path from where it read `SKILL.md`, instead of using `${CLAUDE_SKILL_DIR}`, which only Claude Code expands. `references/browser-state-fixtures.md` uses the same placeholder in its example command.

### code v1.15.0

#### Added
Expand Down
2 changes: 1 addition & 1 deletion plugins/code/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "code",
"description": "Code and planning framework plugin",
"version": "1.15.0",
"version": "1.15.1",
"author": {
"name": "ClosedLoop",
"email": "support@closedloop.ai"
Expand Down
2 changes: 1 addition & 1 deletion plugins/code/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -330,7 +330,7 @@ Staged pipeline for inventorying a Claude Design export into reviewable findings

### `guided-manual-qa`

Derives and runs an interactive, evidence-recorded manual QA session for a code change, ticket, branch, or pull request. Resolves the exact worktree and head under test, maps candidate checkpoints against passing exact-head E2E coverage, and schedules human QA only for the uncovered remainder. Prepares a trustworthy local environment (worktree-owned services, verified origin, proven persistence chain), writes a durable Markdown QA record outside the tracked tree before the first checkpoint, and proves each checkpoint's oracle before presenting it. The human confirms checkpoint by checkpoint with `PASS`, `FAIL`, or `BLOCKED`; agent observations are recorded as supporting evidence, never as human confirmation. Ships a bundled Playwright launcher (`scripts/dist/launch-interactive-browser.mjs`, Node 18+) that opens the interactive browser with preloaded localStorage fixtures and an optional `--ready-selector` gate. Scripts are TypeScript under `tools/guided-manual-qa/src/` with the built bundle committed to `skills/guided-manual-qa/scripts/dist/`. Performs no source changes or external writes without separate authorization.
Derives and runs an interactive, evidence-recorded manual QA session for a code change, ticket, branch, or pull request. Resolves the exact worktree and head under test, maps candidate checkpoints against passing exact-head E2E coverage, and schedules human QA only for the uncovered remainder. Prepares a trustworthy local environment (worktree-owned services, verified origin, proven persistence chain), writes a durable Markdown QA record outside the tracked tree before the first checkpoint, and proves each checkpoint's oracle before presenting it. The human confirms checkpoint by checkpoint with `PASS`, `FAIL`, or `BLOCKED`; agent observations are recorded as supporting evidence, never as human confirmation. Ships a bundled Playwright launcher (`scripts/dist/launch-interactive-browser.mjs`, Node 18+) that opens the interactive browser with preloaded localStorage fixtures and an optional `--ready-selector` gate. Scripts are TypeScript under `tools/guided-manual-qa/src/` with the built bundle committed to `skills/guided-manual-qa/scripts/dist/`. Performs no source changes or external writes without separate authorization. The same skill directory also carries `agents/openai.yaml` display metadata so Codex can load it as a skill.

---

Expand Down
7 changes: 4 additions & 3 deletions plugins/code/skills/guided-manual-qa/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,10 +34,11 @@ State the proposed scope, environment, fixtures, and known gaps before launching
- Distinguish mock-backed and database-backed surfaces explicitly. A fixture-only prototype may need no database, while adjacent production consumers of the same shared component may require a locally seeded app/API stack. Record which checkpoint uses which data source; do not describe prototype fixtures as seeded production data or skip a required production-consumer regression because the prototype renders.
- A prototype is an optional harness for eligible shared code, not a manual-QA destination by itself. Do not add prototype-only E2E tests. Assess repository-supported E2E coverage for affected production code separately; neither a manual finding nor Storybook reachability automatically requires a new E2E test.
- Start services from the resolved worktree. Record the launch commands, working directories, process identifiers, ports, and health checks. Prove that each tested listener belongs to this worktree using the strongest available evidence: process command and cwd, parent process, build or commit marker, service metadata, or a repository-provided diagnostic endpoint.
- If the QA session spans tool calls or worker turns, keep its services under a repository-supported or OS-supported owner that survives that boundary. Record how to inspect and stop that owner. On resume, read the existing QA record and recheck the current head, owned processes, listeners, data target, and exact route; earlier PIDs and ready checks are historical evidence, not proof that the environment is still available.
- After that proof succeeds, launch the UI the user requested for the intentional interactive session. Do not open unrelated surfaces or claim an unlaunched surface was exercised.
- Record feature-flag assignments, roles, permissions, account or fixture identity, and other state that changes the observable result. Redact credentials and secrets.
- When the matrix depends on browser-local state such as feature-flag fixtures, configure it before the first app navigation through a supported browser-context mechanism. Read [references/browser-state-fixtures.md](references/browser-state-fixtures.md). Do not make the human open DevTools or paste JavaScript, and do not use `javascript:` URLs, raw CDP, or the user's ordinary browser profile.
- When the repository has Playwright installed and no stronger repository launcher exists, run the bundled launcher `${CLAUDE_SKILL_DIR}/scripts/dist/launch-interactive-browser.mjs` (Node 18+, no install step) with the bootstrapped repository root as the working directory to open the intentional interactive window with preloaded state. Keep its process alive through the checkpoint and stop only the launcher processes created for the session.
- When the repository has Playwright installed and no stronger repository launcher exists, run the bundled launcher `scripts/dist/launch-interactive-browser.mjs` (Node 18+, no install step) to open the intentional interactive window with preloaded state. Keep its process alive through the checkpoint and stop only the launcher processes created for the session. The launcher path is relative to this skill's directory, not to the repository under test. Run the launcher with the bootstrapped repository root as the working directory, so it resolves Playwright from that repository, and invoke it by the absolute path you resolve from the directory where you read this `SKILL.md`.
- Run only the automated prechecks that make the interactive session meaningful. Follow repository policy for headless browser tests and displayless Electron tests. An explicitly requested interactive manual session may open the requested UI; automated tests must not become visible as a side effect.

Use bounded recovery. Never repeat an unchanged failing launch command. Make at most one targeted repair per documented launch path, capture the exact command and failure, then move to a documented fallback or mark the affected checkpoint `BLOCKED`. Do not improvise an unverified substitute and present it as equivalent.
Expand Down Expand Up @@ -73,8 +74,8 @@ For each checkpoint:

1. Put the application in the required state using safe local setup.
2. Complete and record the oracle proof above.
3. Present exactly one small human action or observation, its expected result, and what evidence to capture.
4. Wait for the human to report `PASS`, `FAIL`, or `BLOCKED`, plus the observed result. Do not advance on an assumption.
3. If the checkpoint asks the human to inspect a UI, open the requested interactive window or app yourself, verify its settled origin and a visible control owned by that route, and keep it available while the human tests. Then present exactly one small human action or observation, its expected result, and what evidence to capture. If the window or route is unready, keep the checkpoint pending and report the setup limitation instead of giving a test.
4. Wait for the human to report `PASS`, `FAIL`, or `BLOCKED`, plus the observed result. After presenting the checkpoint, stop tool calls and do not advance on an assumption; resume only when the human responds or asks for setup help.
5. Re-check disputed expectations before classifying a mismatch, then write the status, exact actual behavior, confirmer, timestamp, evidence location, and any oracle correction to the QA record.
6. Adapt the remaining plan. A failure may require a minimal reproduction, a narrower diagnostic checkpoint, or skipping only dependent checkpoints. A blocked prerequisite must not silently erase the dependent coverage.

Expand Down
4 changes: 4 additions & 0 deletions plugins/code/skills/guided-manual-qa/agents/openai.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
interface:
display_name: "Guided Manual QA"
short_description: "Run evidence-based interactive manual QA"
default_prompt: "Use $guided-manual-qa to derive and guide a manual QA session for this change set."
Original file line number Diff line number Diff line change
Expand Up @@ -31,10 +31,10 @@ const context = await browser.newContext({

Values in Web Storage are strings. Serialize structured fixtures once, before building the entries. Validate the target with `new URL()` and use its exact `.origin`; never accept an arbitrary script or expression as fixture input.

The bundled launcher implements this path without writing browser state into the repository. Run it with the bootstrapped repository root as the working directory, using the absolute launcher path given in SKILL.md:
The bundled launcher implements this path without writing browser state into the repository. Run it with the bootstrapped repository root as the working directory. Its path is relative to the skill directory, so resolve the absolute path as SKILL.md describes:

```bash
node "<launcher path from SKILL.md>" \
node "<absolute skill directory>/scripts/dist/launch-interactive-browser.mjs" \
--url 'http://localhost:3000/path-under-test' \
--storage-file /private/untracked/storage-fixture.json \
--state-name 'flag-matrix-state' \
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@ Copy this template to the chosen durable, untracked QA-record location. This phy
- Browser state origin and non-secret keys:
- Browser state verification:
- Disposable browser context/profile and cleanup path:
- Interactive window or app owner, settled route/control, and last live verification time:
- Feature-flag assignments:
- Role / permissions:
- Local or non-production account / tenant:
Expand All @@ -39,6 +40,8 @@ Copy this template to the chosen durable, untracked QA-record location. This phy
| Service | Launch command | Cwd | PID | Listener / endpoint | Health result | Proof it maps to this worktree |
| --- | --- | --- | --- | --- | --- | --- |

Record the service owner or supervisor that keeps each required process alive across tool calls or worker turns, plus the supported command that inspects and stops only that owner. Recheck these rows after a resume; prior readiness is historical evidence.

### Persistence runtime proof

- Expected repository data-service lane (Docker / Compose / native / other) and evidence:
Expand Down
Loading