Skip to content

fix!: only report a scan as secure when every check was graded - #9

Merged
x1xhlol merged 2 commits into
ZeroLeaks:mainfrom
neoarz:fix/honest-scan-results
Sep 25, 2026
Merged

x1xhlol merged 2 commits into
ZeroLeaks:mainfrom
neoarz:fix/honest-scan-results

Conversation

@neoarz

@neoarz neoarz commented Sep 25, 2026

Copy link
Copy Markdown
Member

A scan could print "SECURE 100/100" and exit 0 without checking the
target. Grader failures counted as refusals, failed probes were dropped,
a mistyped filter ran zero probes, and unset options swapped the target
model. #5 fixed the case where every call fails;
partial failures still passed.

  • Add an "inconclusive" verdict for scans that found nothing but had
    turns or probes error, ran none, or aborted. It scores 0.
  • Report per-mode coverage (graded checks, failed checks with their
    errors, and checks skipped by the time budget or an abort) and the
    model each role used.
  • Let evaluator and judge failures surface instead of guessing from
    keywords, and ignore extracted text on "none" verdicts.
  • Give each injection probe its own conversation, so the two halves of
    dual mode stop clearing each other's history and the injection
    transcript is kept.
  • Stop unset scan options from overwriting the defaults.
  • The CLI exits 0 when secure, 1 when vulnerable, and 2 when there is no
    verdict. Before scanning, it rejects unknown categories and
    severities, bad counts, a --duration too short to run anything, an
    unreadable --file, and an unwritable -o. It flushes stdout before
    exiting so piped --json output isn't cut off.
  • Add an E2E suite (bun test) that runs the CLI against a local mock
    LLM. It needs no API key or network and saves each run to
    test/e2e/artifacts/.
  • Run lint, typecheck, build, and the E2E suite in CI on every pull
    request. No API keys are needed, and the test artifacts are uploaded
    even when a run fails.

BREAKING CHANGE: overallVulnerability and injectionVulnerability can be
"inconclusive", and ScanResult has new required coverage and models
fields. Evaluator.evaluate and InjectionEvaluator.evaluate throw when
the grading model fails. Usage errors and failed scans exit 2 instead
of 1. createTarget defaults to anthropic/claude-sonnet-5 instead of
x-ai/grok-3-mini. runSecurityScan uses the documented defaults for
options left unset (models, inspector, orchestrator) instead of turning
them off.

A scan could print "SECURE 100/100" and exit 0 without checking the
target. Grader failures counted as refusals, failed probes were dropped,
a mistyped filter ran zero probes, and unset options swapped the target
model. ZeroLeaks#5 fixed the case where every call fails;
partial failures still passed.

- Add an "inconclusive" verdict for scans that found nothing but had
  turns or probes error, ran none, or aborted. It scores 0.
- Report per-mode coverage (graded checks, failed checks with their
  errors, and checks skipped by the time budget or an abort) and the
  model each role used.
- Let evaluator and judge failures surface instead of guessing from
  keywords, and ignore extracted text on "none" verdicts.
- Give each injection probe its own conversation, so the two halves of
  dual mode stop clearing each other's history and the injection
  transcript is kept.
- Stop unset scan options from overwriting the defaults.
- The CLI exits 0 when secure, 1 when vulnerable, and 2 when there is no
  verdict. Before scanning, it rejects unknown categories and
  severities, bad counts, a --duration too short to run anything, an
  unreadable --file, and an unwritable -o. It flushes stdout before
  exiting so piped --json output isn't cut off.
- Add an E2E suite (bun test) that runs the CLI against a local mock
  LLM. It needs no API key or network and saves each run to
  test/e2e/artifacts/.
- Run lint, typecheck, build, and the E2E suite in CI on every pull
  request. No API keys are needed, and the test artifacts are uploaded
  even when a run fails.

BREAKING CHANGE: overallVulnerability and injectionVulnerability can be
"inconclusive", and ScanResult has new required coverage and models
fields. Evaluator.evaluate and InjectionEvaluator.evaluate throw when
the grading model fails. Usage errors and failed scans exit 2 instead
of 1. createTarget defaults to anthropic/claude-sonnet-5 instead of
x-ai/grok-3-mini. runSecurityScan uses the documented defaults for
options left unset (models, inspector, orchestrator) instead of turning
them off.
… fail

- A scan callback that threw synchronously skipped its .catch() and
  landed in the scan's own error handling. In an injection scan, probes
  that were already graded were counted again as failed (coverage went
  negative), and three in a row aborted the scan as inconclusive. In an
  extraction scan, a throwing onFinding aborted the scan and skipped the
  leak-status update, so a high finding was reported as medium. Every
  callback now goes through one helper that ignores its failure.
- A report that failed to save after the scan exited with the verdict's
  code, so a secure scan exited 0 with no report written. It now exits 2;
  the verdict is still printed.
- Injection transcript messages were numbered with the extraction turn
  counter, so an injection-only scan numbered every message 0. Each
  message now carries its probe's position.

New E2E scenarios cover each case and fail on the previous code.
Callbacks run through test/e2e/library-scan.ts, which calls
runSecurityScan directly.
@neoarz
neoarz force-pushed the fix/honest-scan-results branch from c18d39e to c4d3a6f Compare September 25, 2026 19:04
@x1xhlol
x1xhlol merged commit 3f36be8 into ZeroLeaks:main Sep 25, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants