Skip to content

Repository files navigation

AssertLedger

Does your regression test actually catch the bug?

AssertLedger runs the same test against fixed code, a known fault and a neutral control. You get a verdict, the observations behind it and an evidence file you can replay.

English · Français

Try the example · Understand the result · Use your repository · Documentation

1.4 · node:test and bun:test · CLI, SDK and MCP · MIT

Install in your repository with Node.js 22.15 or later:

npm install --save-dev assertledger@1.4.0
npx assertledger doctor .

The historical correction demo runs from the installed package. The source example below walks through the evidence step by step.

A passing test can miss the bug

Suppose isEven(2) should return true. A regression accidentally inverts the implementation.

// Both tests pass on the correct implementation.
assert.equal(typeof isEven(2), "boolean"); // Also passes when the answer is wrong.
assert.equal(isEven(2), true);             // Detects this particular regression.

AssertLedger makes that distinction explicit:

Same candidate test Fixed code Known fault Neutral control Result
“Returns a boolean” Pass Pass Pass WEAK_ORACLE for this fault
“Two is even” Pass Assertion failure Pass Eligible for selection
Test crashes or times out — Operational error — No credited detection

The operator supplies the fault and the neutral control. AssertLedger does not invent their meaning. Repeated runs and base tests check that the observed difference can be attributed to the candidate.

flowchart LR
    T[Same candidate test] --> R[Fixed code]
    T --> B[Known fault]
    T --> N[Neutral control]
    R --> E[Recorded observations]
    B --> E
    N --> E
    E --> V[Deterministic verdict]
    V --> M[Replayable evidence]
Loading

Try the example

You need Git, Node.js 22.15+ and pnpm 11.1.2. The example uses the built-in node:test adapter.

git clone --branch v1.4.0 https://github.com/hoklims/assertledger.git
cd assertledger
pnpm install --frozen-lockfile
pnpm build

The bundled example contains a correct parity function, an inverted version, a neutral equivalent, and the two tests above. Inspect its request, then run:

node dist/cli.js demo --allow-unsafe-execution
# Or inspect the complete manifest directly:
node dist/cli.js verify examples/node-test/request.json --allow-unsafe-execution --json

demo copies the shipped example to a disposable temporary directory and removes it afterward. Its result demonstrates the installed AssertLedger package only; it is not evidence about your repository.

Run trusted code only. --allow-unsafe-execution authorizes local code execution. This backend is explicitly UNSANDBOXED. Use a trusted checkout; it cannot contain hostile code.

Expected decision:

{
  "status": "VERIFIED",
  "selectedCandidateIds": ["strong"],
  "reasonCodes": ["POLICY_SATISFIED"]
}

That is the decision section of the full manifest. The weak candidate is marked WEAK_ORACLE. The manifest also records the controls, attempts, observed outcomes, digests and execution limits.

Save and replay the evidence

Create demo.mjs at the repository root with the following content. This writes UTF-8 consistently on Windows and Linux and uses the same example through the SDK:

import { readFile, writeFile } from "node:fs/promises";
import { AssertLedger } from "./dist/index.js";

const ledger = new AssertLedger();
const request = JSON.parse(await readFile("examples/node-test/request.json", "utf8"));
// Review the trusted example before explicitly authorizing its execution.
request.isolation.acknowledgedUnsafeExecution = true;
const manifest = await ledger.verify(request);
await writeFile("manifest.json", JSON.stringify(manifest, null, 2), "utf8");
console.log(manifest.decision);
node demo.mjs
node dist/cli.js replay manifest.json --json

All five replay fields should be true: valid, schemaValid, decisionDigestValid, artifactDigestValid and decisionSemanticsValid. Replay requires no model and does not run tests again.

Understand the result

Campaign verdict What it tells you Next step
VERIFIED Selected tests meet the declared policy for these worlds and attempts. Review the fault, controls and evidence before accepting the test.
REJECTED No candidate meets the declared policy. Read candidate reasons; strengthen the assertion or correct the declared worlds.
INCONCLUSIVE The observations do not support a stable decision. Inspect unstable runs, discovery and operational errors.
ENGINE_ERROR The campaign could not produce a usable result. Fix the environment or configuration, then rerun.

A timeout, syntax error, collection failure or process crash never counts as a detected bug. A rejection concerns the declared fault model; the test may still have value elsewhere.

Replay checks integrity and decision consistency. It does not authenticate whoever produced the observations, prove general program correctness or guarantee permanent freedom from flaky tests.

Use your repository

Start with a static diagnostic. It reads the repository without running its tests or writing files:

node dist/cli.js doctor path/to/your-repository
node dist/cli.js doctor path/to/your-repository --json

To preview initialization and a read-only agent connection as one conflict-checked operation:

node dist/cli.js setup path/to/your-repository --client codex --dry-run
node dist/cli.js setup path/to/your-repository --client codex --write

Use --client claude-code for Claude Code. The preview is also the default when neither mode flag is present. Setup checks every initialization and connection target before its first managed-file write. If a connection conflict appears after initialization, setup removes only init files that this invocation created and that still match its exact bytes. Anything it cannot safely restore is reported as PARTIAL_FAILURE with exit code 5 and an explicit unresolved-file list.

WOULD_CREATE means a configuration can be planned. You still supply the candidate and controls. The initialization guide explains init, the configuration and evidence lock. Detection of a framework is not proof that AssertLedger can execute it.

After initialization, runtime doctor can check Node, the reporter, discovery and assertion attribution with doctor --runtime --allow-unsafe-execution. For a refusal, use explain CODE to get a safe next action.

Bun 1.4.2

For a bun:test repository, run assertledger doctor . --framework bun:test. After init and the runtime doctor, author v3 regression candidates using assertSame from assertledger/bun. The Bun adapter pins version 1.4.2 and attributes failures from this helper. Native Bun expect failures remain operational failures and cannot count as target detection. See the Bun migration guide for the request, execution and replay contract. The execution mode is explicitly unsandboxed and intended for trusted local code.

Qualify a committed regression test

Choose the buggy commit (BEFORE), its correction (AFTER) and a neutral control (NEUTRAL). The candidate comes from AFTER; the exact same bytes run in all three worlds. Replace the paths, revisions and neutral reason below with your own:

node dist/cli.js check path/to/your-repository --before BEFORE --after AFTER --neutral NEUTRAL --neutral-reason "Explain why this control preserves the expected behavior" --test tests/regression.test.js --base-test tests/base.test.js --out .assertledger/evidence-001 --allow-unsafe-execution
Required for this first Git workflow Why
Committed JavaScript node:test candidate Each world receives the same recorded test.
No declared runtime dependencies This workflow does not install or transport dependencies.
An unchanged base test in every revision The control must not change with the correction.
The same file paths, apart from the candidate File additions, deletions and renames are not qualified yet.

The command saves summary.md, executed-request.json and manifest.json in a new output directory. manifest.json is published last; its presence marks a complete result. The saved request refers to a temporary snapshot that has been removed; preserve the Git revisions if you need to run again. Read the Git workflow guide for limits and neutral-control semantics.

Use it with an agent

An agent can propose a candidate; the deterministic engine evaluates the observations. Use the agent skill, TypeScript SDK or MCP reference.

Generate a project-scoped Codex configuration from the built CLI:

node dist/cli.js connect path/to/your-repository --client codex

This previews the configuration and packaged skill. Add --write to install both; different existing content is preserved and reported as a conflict. The generated MCP server starts read-only. Candidate execution requires a separate explicit opt-in. See developer entry points for project trust and reload requirements.

Use --client claude-code for Claude Code or --client mcp for a generic descriptor. disconnect --client codex --write removes only byte-identical owned files. The client guide covers installation and removal.

Documentation

You want to… Start here
Set up a repository Initialization · Static audit
Verify a Bun regression Bun migration
Diagnose a blockage Runtime doctor · Reason-code guidance
Try a historical correction Unicode-regexp example
Qualify a correction or connect Codex Git workflow · Developer entry points
Understand attribution, controls and digests Proof model
Integrate the CLI, SDK or MCP Integration reference
Verify the actual distributed package Distribution checks · CI evidence
Extend an adapter Adapter protocol · Architecture
Review the 1.0 scope Release criteria · Roadmap
Explore advanced evaluation work Profiles · Benchmarks · Calibration

Scientific profiles and calibration retain their own evidence requirements. Their presence does not establish an improvement on an external product or a measured product-market fit.

Contribute

pnpm check
pnpm run smoke:package

Add a failing behavioral test for public contract changes. Keep execution in the engine and decision logic in the pure core. Read CONTRIBUTING.md and SECURITY.md.

Compatibility: AssertLedger is the current name. Legacy TestForge aliases and versioned wire identifiers remain available so existing integrations and evidence can be replayed. See the migration guide.

Licensed under MIT.

About

Verify that regression tests catch the bug they claim to prevent. Deterministic, replayable evidence for agent-assisted development.

Topics

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages