Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 16 additions & 12 deletions .agents/prompts/reviewer.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,18 +4,22 @@ You did not implement this change. Review only. You run with read-only permissio

Read `AGENTS.md`, the active plan, relevant architecture/ADRs, and `docs/evaluation.md`. The calling script supplies the GitHub Issue and the complete feature-branch diff against its base, including uncommitted and untracked working-tree changes; review that complete diff and read repository files for context.

Report findings as **Critical**, **Major**, **Minor**, or **Suggestion**. For each finding include evidence/file location, why it matters, and a recommended action. Check correctness, issue/plan compliance, architecture, security, edge cases, test coverage, reliability, and unnecessary complexity.
Check correctness, issue/plan compliance, architecture, security, edge cases, test coverage, reliability, and unnecessary complexity.

Place every finding under its matching `## Critical`, `## Major`, `## Minor`,
or `## Suggestions` section and give it a stable heading such as
`### C1. Title`, `### M1. Title`, `### Minor 1. Title`, or `### S1. Title`.
Write `None.` when a section has no findings. These identifiers are preserved
by review triage and follow-up Issues.
## Result

Return the complete review as your final message. It must contain the sections
`## Critical`, `## Major`, `## Minor`, `## Suggestions`, and `## Verdict`, in
that order. The calling script stores the review; do not write it to a file.
Return the review as JSON that matches the schema supplied by the calling script (`.agents/schemas/review.schema.json`). Do not write it to a file and do not add text around it.

Under `## Verdict`, write exactly one of `PASS`, `PASS WITH MINOR FINDINGS`, or
`CHANGES REQUIRED` on its own line. Base the verdict on the findings. State
anything you could not verify below the verdict line.
- `findings`: one entry per finding; an empty array when there are none.
- `severity`: `critical`, `major`, `minor`, or `suggestion`.
- `title`: a short, specific name for the finding.
- `evidence`: file locations and what you observed there.
- `impact`: why it matters.
- `recommendation`: the recommended action.
- `verdict`: derived from the findings only.
- `PASS`: no findings.
- `PASS_WITH_MINOR_FINDINGS`: only minor or suggestion findings.
- `CHANGES_REQUIRED`: at least one critical or major finding.
- `limitations`: anything you could not verify, such as checks you could not run. Use an empty string when there is nothing to report. A limitation is not a finding and does not change the verdict.

Do not number the findings; the calling script assigns the identifiers that review triage and follow-up Issues use. A result that breaks these rules is rejected.
4 changes: 2 additions & 2 deletions .agents/prompts/triage-implementer.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,8 +10,8 @@ Read:
- `AGENTS.md`
- the originating GitHub Issue
- the matching active feature plan, when one exists
- the source independent-review artifact
- the approved triage artifact
- the source independent-review artifact (JSON)
- the approved triage artifact (JSON)
- relevant architecture documentation and accepted ADRs
- the current working-tree diff

Expand Down
31 changes: 15 additions & 16 deletions .agents/prompts/triage-reviewer.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,7 @@ implementing fixes and you must not modify repository files.
Read:

- `AGENTS.md`
- the complete source review artifact supplied by the caller
- the findings manifest supplied by the caller
- the review findings supplied by the caller as JSON
- the originating GitHub Issue, when identified
- the matching active feature plan, when one exists
- relevant architecture documentation and accepted ADRs when needed
Expand All @@ -19,7 +18,7 @@ silently downgrade findings.

## Decisions

Classify every manifest entry exactly once:
Classify every finding exactly once:

- `FIX_NOW`: blocks the current feature or is a clear, local, valuable fix.
Critical and Major findings must use this decision.
Expand All @@ -39,22 +38,22 @@ For `DEFER`, also propose:
Do not create GitHub Issues. The calling script owns the human approval gate
and all approved side effects.

## Output format
## Result

Return only tab-separated records, one per manifest entry, in the same order:
Return JSON that matches the schema supplied by the calling script
(`.agents/schemas/triage.schema.json`). Do not write it to a file and do not
add text around it.

```text
<key> <decision> <rationale> <issue-title-or-dash> <recommended-action-or-dash> <acceptance-criteria-or-dash>
```
- `decisions`: one entry per finding of the review, and no others.
- `finding_id`: the `id` of the finding, exactly as supplied.
- `decision`: `FIX_NOW`, `DEFER`, or `ACCEPT`.
- `rationale`: the reason for the decision.
- `followup`: `null` unless the decision is `DEFER`. For `DEFER`, an object
with `title`, `recommended_action`, and at least one entry in
`acceptance_criteria`.

Rules:

- Use exactly six fields separated by literal tab characters.
- Keep every field on one physical line and do not place tabs inside fields.
- Use only `FIX_NOW`, `DEFER`, or `ACCEPT` for the decision.
- Use `-` for the final three fields unless the decision is `DEFER`.
- Do not add Markdown fences, headings, commentary, summaries, or blank lines.
- If the findings manifest is empty, output exactly `NO_FINDINGS`.
A result that omits a finding, decides one twice, refers to an unknown finding,
or breaks the rules above is rejected.

## Boundaries

Expand Down
30 changes: 30 additions & 0 deletions .agents/schemas/review.schema.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
{
"type": "object",
"additionalProperties": false,
"required": ["verdict", "findings", "limitations"],
"properties": {
"verdict": {
"type": "string",
"enum": ["PASS", "PASS_WITH_MINOR_FINDINGS", "CHANGES_REQUIRED"]
},
"findings": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["severity", "title", "evidence", "impact", "recommendation"],
"properties": {
"severity": {
"type": "string",
"enum": ["critical", "major", "minor", "suggestion"]
},
"title": { "type": "string" },
"evidence": { "type": "string" },
"impact": { "type": "string" },
"recommendation": { "type": "string" }
}
}
},
"limitations": { "type": "string" }
}
}
41 changes: 41 additions & 0 deletions .agents/schemas/triage.schema.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
{
"type": "object",
"additionalProperties": false,
"required": ["decisions"],
"properties": {
"decisions": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["finding_id", "decision", "rationale", "followup"],
"properties": {
"finding_id": { "type": "string" },
"decision": {
"type": "string",
"enum": ["FIX_NOW", "DEFER", "ACCEPT"]
},
"rationale": { "type": "string" },
"followup": {
"anyOf": [
{ "type": "null" },
{
"type": "object",
"additionalProperties": false,
"required": ["title", "recommended_action", "acceptance_criteria"],
"properties": {
"title": { "type": "string" },
"recommended_action": { "type": "string" },
"acceptance_criteria": {
"type": "array",
"items": { "type": "string" }
}
}
}
]
}
}
}
}
}
}
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,14 +44,14 @@ A lightweight, model-agnostic repository template for agentic software engineeri
./scripts/verify.sh
./scripts/review-feature.sh 12
./scripts/triage-review.sh \
.agents/reviews/feature-12-player-movement-review-01.md
.agents/reviews/feature-12-player-movement-review-01.json
```

8. Approve the proposed triage and apply its `FIX_NOW` scope:

```bash
./scripts/apply-triage.sh \
.agents/triage/feature-12-player-movement-review-01-triage.md
.agents/triage/feature-12-player-movement-review-01-triage.json
```

The script starts a write-capable agent only after confirmation and verifies
Expand Down
7 changes: 5 additions & 2 deletions docs/agentic-workflow.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,8 +30,11 @@ See `docs/development.md` for the concrete commands.
- ADRs: why significant architecture decisions were made.
- `.agents/plans/`: active implementation state for complex work.
- `.agents/handoffs/`: compressed continuation context.
- `.agents/reviews/`: temporary independent-review artifacts.
- `.agents/triage/`: approved finding decisions and deferred-Issue traceability.
- `.agents/reviews/`: independent-review results as validated JSON, each with a
generated Markdown report.
- `.agents/triage/`: approved finding decisions and deferred-Issue traceability
as validated JSON, each with a generated Markdown report.
- `.agents/schemas/`: the schemas of those results.
- `.agents/lessons/`: recurring failure lessons awaiting/promoting durable rules.
- Git history: what actually changed.
- PR + CI: review discussion and deterministic evidence.
Expand Down
82 changes: 63 additions & 19 deletions docs/development.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,9 +78,13 @@ Agents run with one of two permission profiles:
- `write`: an interactive session that may modify its worktree. Codex runs in
a workspace-write sandbox without approval prompts; Claude accepts edits
automatically. Used for planning, implementation, and applying triage.
- `read-only`: a non-interactive session that cannot modify files. Codex runs
in a read-only sandbox; Claude is limited to its read tools. The script
stores the agent's final message. Used for review and triage.
- `read-only`: a non-interactive session that cannot modify files and gets no
MCP servers, apps, or other tools from the user's configuration. Codex runs
in a read-only sandbox without the user's `config.toml`, with apps, browser
use, computer use, and web search disabled. Claude runs restricted: without user, project, or MCP
configuration, with only its Read, Glob, and Grep tools, and without asking
for any further permission. The script stores the agent's result. Used for
review and triage.

A profile that a provider cannot enforce is an error; an agent is never started
with broader permissions instead.
Expand Down Expand Up @@ -208,31 +212,33 @@ It uses role `reviewer` with the `read-only` profile. The review is
non-interactive: the reviewer cannot modify files and has no network access, so
the script supplies the GitHub Issue and the complete diff against the base
branch, including uncommitted and untracked changes. The reviewer returns its
report and the script stores it. A report without the required sections or a
valid verdict is rejected, and a reviewer that changed the working tree or
result as JSON and the script validates and stores it. An invalid result is
retried once and then rejected, and a reviewer that changed the working tree or
created a commit is reported as an error; in both cases no review is stored.

The review script must run from the matching feature worktree and needs an
authenticated GitHub CLI. It writes numbered artifacts without overwriting
earlier reviews:
authenticated GitHub CLI. Each round writes a numbered pair of files without
overwriting earlier reviews:

```text
.agents/reviews/feature-12-player-movement-review-01.json
.agents/reviews/feature-12-player-movement-review-01.md
.agents/reviews/feature-12-player-movement-review-02.md
```

See "Review and triage data" below for the two files.

Triage an explicit review artifact with an agent independent from the
implementation:

```bash
./scripts/triage-review.sh \
.agents/reviews/feature-12-player-movement-review-01.md
.agents/reviews/feature-12-player-movement-review-01.json
```

The full interface is:

```text
./scripts/triage-review.sh <review-file> [--agent <agent>] [--model <model>]
./scripts/triage-review.sh <review-json> [--agent <agent>] [--model <model>]
```

The triage agent (role `triage`) classifies every finding as:
Expand All @@ -242,44 +248,53 @@ The triage agent (role `triage`) classifies every finding as:
- `DEFER`: valid non-blocking work proposed as a separate follow-up Issue.
- `ACCEPT`: consciously take no action, with an explicit rationale.

The script validates the decisions before showing them: every finding is
decided exactly once, Critical and Major findings are `FIX_NOW`, and a deferred
finding has a follow-up title, action, and acceptance criteria. Invalid
decisions are retried once and then rejected. A review without findings needs
no triage agent.

The script displays the complete proposal before side effects. Only after
interactive approval does it create one GitHub Issue per `DEFER` finding and
write a persistent, uniquely named artifact such as:

```text
.agents/triage/feature-12-player-movement-review-01-triage.json
.agents/triage/feature-12-player-movement-review-01-triage.md
```

That artifact maps the source review findings to their decisions and any
created Issue numbers. Declining the proposal creates neither an artifact nor
Issues. The source review remains unchanged. A separate `create-followups.sh`
is therefore not needed.
Issues. The source review remains unchanged. The artifact is stored only after
every follow-up Issue exists; if creating one fails, nothing is stored and a
new triage reuses the Issues created so far, which it finds by their trace
token. A re-run reuses a follow-up Issue that already exists for a finding.

Deferred follow-up Issue titles include deterministic provenance:

```text
[F02][R01][S9] Concise follow-up title
[F02][R01][S2] Concise follow-up title
```

The script obtains the feature ID from the source Issue title, the review round
from the review filename, and the finding ID from the review. When no feature
ID is available, it falls back to the source Issue number:
The script obtains the feature ID from the source Issue title and the review
round and finding ID from the review. When no feature ID is available, it falls
back to the source Issue number:

```text
[#12][R01][S9] Concise follow-up title
[#12][R01][S2] Concise follow-up title
```

Apply the approved `FIX_NOW` set from the same feature worktree:

```bash
./scripts/apply-triage.sh \
.agents/triage/feature-12-player-movement-review-01-triage.md
.agents/triage/feature-12-player-movement-review-01-triage.json
```

The full interface is:

```text
./scripts/apply-triage.sh <triage-file> [--agent <agent>] [--model <model>]
./scripts/apply-triage.sh <triage-json> [--agent <agent>] [--model <model>]
```

The helper uses role `triage-implementer`. It validates the approved artifact and source review, shows the exact
Expand All @@ -292,6 +307,35 @@ The helper does not commit, push, merge, deploy, or create/close Issues. Inspect
the resulting diff and run another independent review and triage when fixes
require confirmation.

### Review and triage data

Results that an agent writes and a script consumes are JSON. Each review and
each triage is stored as two files with the same name:

- `.json` is the source of truth. The scripts read only this file.
- `.md` is a report generated from the JSON for reading. It is never parsed;
editing it has no effect.

The agent's result must match a schema in `.agents/schemas/`
(`review.schema.json`, `triage.schema.json`). The schema is passed to the
provider CLI and the script checks the result again with `jq`, including rules
a schema cannot express. For a review, the verdict must follow from the
findings: `PASS` without findings, `PASS_WITH_MINOR_FINDINGS` with only minor
or suggestion findings, and `CHANGES_REQUIRED` with at least one critical or
major finding. What a reviewer could not verify belongs in `limitations` and
does not change the verdict.

The script, not the agent, numbers the findings: `C1` (critical), `M1` (major),
`MIN1` (minor), and `S1` (suggestion). Triage decisions and follow-up Issues
refer to those identifiers. When a result is invalid, the agent is asked once
more, with the reasons for the rejection.

Stored artifacts are checked with the same rules as agent results, plus the
fields the scripts own, so editing a stored artifact cannot weaken a decision.

Documents that people maintain, such as the roadmap, requirements, and
architecture, keep Markdown as their source.

Verify again after review fixes, then commit and push:

```bash
Expand Down
2 changes: 1 addition & 1 deletion docs/evaluation.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ decline the proposed decisions.
- **Suggestion:** classify explicitly; it may be deferred when worthwhile or
accepted with a rationale.

The approved `.agents/triage/` artifact is the source of truth for these
The approved `.agents/triage/*.json` artifact is the source of truth for these
decisions. Use `./scripts/apply-triage.sh` to hand only its `FIX_NOW` scope to a
write-capable implementation agent. The helper verifies the result but does not
commit it. Inspect the diff and run a new independent review/triage when needed
Expand Down
Loading
Loading