[FIX]: Evaluate attack verdicts over final traces - #150
Draft
spencrr wants to merge 6 commits into
Draft
Conversation
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Moves XPIA verdict evaluation to the terminal trace and separates it from online stopping.
stop_whenaccepts an explicit evaluator,None, or the default"auto". Auto mode reuses the verdict evaluator only for framework-owned conditions whose detected result is stable as turns are appended, such as cumulative tool and side-effect checks.An identical stop/verdict evaluator is not called again when its latest result already covers the terminal trace. Final evaluation remains inside the active session and injection stack, observability downgrades are recorded without mutating response metadata, and cleanup failures discard otherwise successful verdict evidence and return
ERROR.Depends on #149. Because the branches live on a fork, this PR temporarily includes lower-layer diffs and targets
main; those diffs disappear as dependencies merge.Breaking changes
Behavioral change: unknown or stochastic verdict evaluators no longer run on every prefix by default. They evaluate once over the terminal trace and may run to
max_turnsunless an explicitstop_whenis supplied. Auto-stopped attacks expose online evaluator feedback to adaptive drivers.Checklist
pre-commit run --all-filespassesValidation: 116 cross-layer attack/probe/runner tests and 726 broad unit tests pass, with two known baseline cases deselected. Strict documentation build, all-files pre-commit, static checks, and privacy scan pass.