Use validated JSON for review and triage data - #39
Merged
Merged
Conversation
Review results and triage decisions are now JSON that must match a schema in .agents/schemas/. The schema is passed to the provider CLI and the scripts check the result again with jq through one shared library, scripts/lib/review-data.sh, which also renders the Markdown reports. Stored artifacts are checked with the same rules as agent output. Invalid output is retried once and then rejected before any side effect. The scripts assign the finding identifiers, a verdict must follow from its findings, and a triage is complete only when every deferred finding has its follow-up Issue. The duplicated AWK finding parsers are gone, and the review, triage, and apply-triage tests run in isolated repositories. Closes #37 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Real Claude and Codex review sessions showed that read-only agents still received MCP servers, apps, and other remote tools from the user's configuration. Claude read-only sessions now run restricted with only Read, Glob, and Grep; Codex read-only sessions ignore the user's config.toml and disable apps, browser use, computer use, and web search. Also from the reviews: stored artifacts constrain script-owned fields, including finding identifiers, Issue references, and copied review data; a triage is stored only after every follow-up Issue exists; the retry tells the agent why its result was rejected; Claude's exit status is returned; and every validation rule has a test. Refs #37 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
.agents/schemas/review.schema.jsonandtriage.schema.jsondefine the result an agent must return.scripts/lib/review-data.shvalidates agent results and stored artifacts with the same rules and renders the Markdown reports. The launcher passes the schema to the provider (codex exec --output-schema,claude --json-schema)..jsonis the source of truth;.mdis generated for reading and never parsed.FIX_NOW, and a deferred finding has a complete follow-up.C1,M1,MIN1,S1) and builds stored artifacts from allowed fields only.apply-triage.sh, and a new triage round reuses Issues already created.triage-review.shandapply-triage.shtake the.jsonartifact.Issue / acceptance criteria
Closes #37
Risk
Verification evidence
./scripts/verify.shpassed (9 checks)Exercised against a real Codex session (
gpt-6-astra): review with schema, storage and rendering, and triage with schema up to the approval question (declined, so no Issues were created).Independent review, two rounds with Codex through
review-feature.sh:followupfield was accepted (Minor); the approval preview omitted the follow-up action and criteria (Minor). All fixed, with tests. Not reviewed again after these fixes.Not verified: the Claude side of structured output. The
claudeCLI on this machine is not authenticated, so how Claude returns the structured result (structured_outputin the JSON envelope) is assumed and covered with fake agents only.Update: real Claude run and agent isolation
After the Claude CLI was logged in, the Claude path was exercised for real: review and triage with schema, through
structured_outputas assumed.fable), CHANGES REQUIRED with 9 findings. All addressed: tests for every validation rule;jqfound by the tests wherever it is installed; triage stored only after every follow-up Issue exists; script-owned fields constrained (finding ids such as../S1, positive integer Issue and round, copied review data); Claude's real exit status returned and the structured-output path tested; isolation of read-only sessions (below); a pre-JSON review is no longer a re-review input; the retry tells the agent why its result was rejected; consistent argument order and deterministic test names.--restricted --strict-mcp-config --permission-mode dontAsk --tools Read,Glob,Grep. Codex now runs with--ignore-user-configand apps, browser use, computer use, and web search disabled. Both were re-probed: they read files, cannot write, and get no MCP tools.Agent involvement
Planner: —
Implementer: Claude (Opus 5.5)
Reviewer: Codex (gpt-6-astra)
Production impact
None.
🤖 Generated with Claude Code