Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
58 commits
Select commit Hold shift + click to select a range
7981988
docs: freeze all 83 completed audit task identities (refs #131)
flyingrobots Oct 1, 2026
dd31a24
docs: audit first five remaining completed tasks (refs #131)
flyingrobots Oct 2, 2026
7fe7513
docs: prevent issue reference from parsing as heading (#131)
flyingrobots Oct 2, 2026
cd2b08d
docs: audit identity layers and chunking criteria (refs #131)
flyingrobots Oct 2, 2026
e918907
docs: audit four completed flat-layout tasks (refs #131)
flyingrobots Oct 2, 2026
41d7459
docs: audit reference tasks with isolated validation (refs #131)
flyingrobots Oct 2, 2026
99a27b7
docs: audit conformance-oracle acceptance criteria (refs #131)
flyingrobots Oct 2, 2026
d1d498b
docs(audit): record benchmark admission gap (#131, #142)
flyingrobots Oct 2, 2026
830e3f7
docs(audit): record missing filename enforcement (#131, #144)
flyingrobots Oct 2, 2026
0b83963
docs(audit): verify durable protocol ports and failure fakes (#131)
flyingrobots Oct 2, 2026
f2732fc
docs: audit immutable segment criteria and sealing gap (#131, #146)
flyingrobots Oct 2, 2026
225eab3
Audit: distinguish merged corrections from remaining criteria (#131)
flyingrobots Oct 2, 2026
4f45dbf
Audit: assign promised filesystem integration evidence (#147)
flyingrobots Oct 2, 2026
bb8eb4a
Audit: record reproduced catalog platform-admission bypass (#150)
flyingrobots Oct 2, 2026
cb2c185
Docs: record merged admission and catalog evidence fixes (#131)
flyingrobots Oct 2, 2026
ad3490a
Docs: avoid heading ambiguity in audit issue reference (#131)
flyingrobots Oct 2, 2026
95c733c
docs: record catalog model evidence gap (#131, #166)
flyingrobots Oct 3, 2026
700c7bf
docs: audit catalog codec and scan-cost acceptance (#131, #168)
flyingrobots Oct 3, 2026
8552dcd
docs: audit initialization authority evidence (#131)
flyingrobots Oct 3, 2026
aab0bb0
docs: record corrupt partial-seal discard finding (#131)
flyingrobots Oct 3, 2026
404c15e
Audit: record duplicate refusal bypass in truncated recovery (#173)
flyingrobots Oct 3, 2026
fecf210
docs: record reviewed PR landing progress (#131 #132)
flyingrobots Oct 3, 2026
17ce2d4
Docs: record three reviewed landings and stacked dependency (#131 #132)
flyingrobots Oct 3, 2026
a097db7
Docs: record verified public stage landing (#131 #147)
flyingrobots Oct 3, 2026
219f6ed
docs: record verified #159 landing and #160 review start (#131)
flyingrobots Oct 3, 2026
653d848
docs: record verified #160 landing and #159 mainline checks (#131)
flyingrobots Oct 3, 2026
99cd488
docs: retain #161 landing blockers and green #160 integration (#131)
flyingrobots Oct 3, 2026
a61b0c2
Docs: record migration closure and reader-fence landing blocker (#131)
flyingrobots Oct 3, 2026
03265ed
Docs: record reader-fence approval and infrastructure gate (#131)
flyingrobots Oct 3, 2026
886cdfc
Docs: record reader-fence merge and migration integration (#131)
flyingrobots Oct 3, 2026
996a143
Docs: record approved migration merge and next landing candidate (#131)
flyingrobots Oct 3, 2026
cfda0c9
Docs: record migration compatibility landing evidence (#131)
flyingrobots Oct 3, 2026
75a81c3
Docs: record reviewed migration compatibility merge (#131)
flyingrobots Oct 3, 2026
dadf621
Docs: record catalog model landing validation (#131)
flyingrobots Oct 3, 2026
b4ba6a1
docs: record #167 landing and #170 integration under #131
flyingrobots Oct 3, 2026
598a0f8
docs: close #167 and #170 landing evidence under #131
flyingrobots Oct 3, 2026
72cd23b
docs: record #164 integration and landing gates under #131
flyingrobots Oct 3, 2026
ca79d95
docs: reconcile late #164 and #165 landing findings under #131
flyingrobots Oct 3, 2026
a235619
docs: record #164 merge and #165 integration under #131
flyingrobots Oct 3, 2026
d617678
docs: record verified #165 landing under #131
flyingrobots Oct 3, 2026
2a31df2
docs: record #165 mainline validation and #163 review progress (#131)
flyingrobots Oct 3, 2026
5f372e5
Docs: record final #163 review and dependency landing preparation (#131)
flyingrobots Oct 3, 2026
ca43b93
Docs: record #163 reviewed mainline integration (#131)
flyingrobots Oct 3, 2026
420405c
Docs: record #106 landing and #104 reviewed candidate (#131)
flyingrobots Oct 4, 2026
c64fa5a
Docs: record dependency landings and mainline contention blocker (#131)
flyingrobots Oct 4, 2026
2a23c46
Docs: record isolated locator correction and final checks (#178)
flyingrobots Oct 4, 2026
56becd2
Docs: record integrated Rustix checks and BLAKE3 landing gates (#131)
flyingrobots Oct 4, 2026
6878b77
Docs: record capability candidate and hosted benchmark evidence run (…
flyingrobots Oct 4, 2026
c91692b
Docs: retain same-host BLAKE3 benchmark measurements (#94)
flyingrobots Oct 4, 2026
6a86622
Docs: record green capability checks and remaining M4 integration (#131)
flyingrobots Oct 4, 2026
0c2d1ff
docs: close independent benchmark evidence review (#94)
flyingrobots Oct 4, 2026
e9b34eb
docs: record scoped GC fuzz closure and remaining PR107 queue
flyingrobots Oct 4, 2026
fd97904
docs: record reviewed reader-lock coordinate closure (#107)
flyingrobots Oct 4, 2026
cc3ff30
docs: record GC provenance and receipt review closure (#107)
flyingrobots Oct 4, 2026
4e58eb5
docs(#131): record closure of original #107 inline findings
flyingrobots Oct 4, 2026
66e276b
docs: record remaining PR #107 integration obligations
flyingrobots Oct 4, 2026
08ebb5d
docs: record focused #107 integration runtime evidence
flyingrobots Oct 4, 2026
3aed1c8
docs: normalize #107 receipt terminal whitespace
flyingrobots Oct 4, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
96 changes: 96 additions & 0 deletions docs/audits/completed-roadmap-criteria/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
# Completed-roadmap audit scope

This directory continues issue #131. The criterion-by-criterion audit is in
progress; the scope manifest is not a completion verdict.

The original roadmap at commit
`1a586d83d5750083172d440f90e7b786d540ff0e` has 83 checked task entries.
The first-pass roadmap at commit
`4b9c38930f988911ab020b7c42e9221b721933af` leaves 64 checked and reopens 19.
`scope.tsv` records every original task, its original line, and its first-pass
state. Originally unchecked tasks and feature-level checkboxes are excluded.

The [foundation verdicts](foundations.md),
[identity-layer and chunking verdicts](identity-layers-and-chunking.md), and
[flat-layout verdicts](flat-layout.md), and
[reference-store and read verdicts](reference-store-and-reads.md), and
[conformance-oracle verdicts](conformance-oracles.md), and
[benchmark-baseline verdict](benchmark-baseline.md), and
[architecture verdicts](architecture.md), and
[immutable-segment verdicts](immutable-segments.md) cover the
first 33 remaining checked tasks. Each separates acceptance and mainline
delivery and names inspected evidence and limits. The individual verdicts
retain their inspected historical coordinates. T-06.3 has unresolved
acceptance scope. T-09.1's canonical report-admission gap was owned by
issue #142 and has since been delivered as recorded below. The historical T-10.2
forbidden-filename gap was owned by issue #144.
T-11.2 exposes writable stage authority after sealing; issue #146 owns that
correction. T-11.3 still lacks its originally named integration-test artifact;
the living-documentation correction in #69 does not fulfill that requirement.

The [catalog publication verdict](catalog-publication.md) adds T-12.2 at
main `b50dbd4cb4cee286aea1aa0352152a232197dda1`: a public repository-task
constructor bypasses production platform admission, reproduced independently
and owned by #150. T-12.3 now has a separate verdict at main `6051abb25a9fd33ae7ee0de5614514b709a4d82a`: restart examples pass, but generated independent catalog-model evidence is missing; issue #166 owns that correction. T-12.1 now has a separate functional/codecs verdict on the same main revision, with its source-text scan-cost evidence gap owned by #168. The [store initialization verdict](store-initialization.md) adds T-13.1 on that same main revision: the authority-producing path correctly propagates lock identity refusal, but its cited runtime evidence survives ignoring that refusal; issue #169 owns the acquisition-boundary evidence repair. Thirty-seven remaining checked tasks now have recorded verdicts; twenty-seven other checked tasks and nineteen reopened entries still need full accounting.

## Mainline delivery after the inspected snapshot

PR #143 delivered canonical benchmark report admission as
`b50dbd4cb4cee286aea1aa0352152a232197dda1`. The
[issue #142 receipt](https://github.com/flyingrobots/keep/issues/142#issuecomment-5946020879)
records exact-head validation, bounded parser fuzzing and forty-five
mainline task/library laws in each build mode. This resolves the identified
admission gap; it does not assert a fresh performance measurement or replace
the historical T-09.1 baseline evidence.

PR #149 delivered the two catalog evidence-anchor corrections as
`82374a995df095106aefe52f87ab3cb26184639d`. The
[current-head merge gate](https://github.com/flyingrobots/keep/pull/149#issuecomment-5946195690)
records the corrected independent review, green required hosted checks and
fresh Docker execution of the two ordering and sixteen publication fixture
laws in both modes on the byte-identical target runtime. Issue #148 is
closed. These documentation corrections neither resolve T-12.2's platform
admission bypass nor close #150.

PR #135 delivered the T-06.4 capacity-bounded reference-store memory contract
as `07bf0b8f4315305e248e1628b339243e5e59ee0b`. The
[issue #74 receipt](https://github.com/flyingrobots/keep/issues/74#issuecomment-5945001183)
records exact-head validation and debug/release mainline allocation laws.
This does not claim constant total memory or a successful four-GiB ingestion.

PR #145 delivered forbidden Rust source-filename enforcement as
`88f35c417abeb62ef72c65b3a1904151cc2e3ee9`. The
[issue #144 receipt](https://github.com/flyingrobots/keep/issues/144#issuecomment-5944930436)
records mainline debug/release policy laws and exact-head validation.

PR #134 delivered single authentication per selected layout occurrence as
`8d902516e682361882bc5c9902de296ce5c9de85`. The
[issue #71 receipt](https://github.com/flyingrobots/keep/issues/71#issuecomment-5945833591)
records debug/release mainline authentication and accounting laws. T-06.3's
immediate-output criterion remains unmet; issue #71 stays open. No original
criterion or checkbox has been changed.

PR #136 delivered current v1 format documentation as
`99551ece786d47ef62ecff785a2e24261a911e85`. Its corrected ledger names the
actual filesystem unit laws. T-11.3's original named
`tests/segment_filesystem_stage.rs` integration target remains absent, so
closing #69 does not close this remaining audit gap. Executable
[issue #147](https://github.com/flyingrobots/keep/issues/147) owns the named
public integration target and its mainline evidence under the #131 audit.

Validation now uses separate Docker build directories per source clone.
A shared target directory reused a stale test binary across clones; those
earlier runs do not establish source-specific validation. Fresh isolated
debug/release keep suites, formatting, workspace Clippy, and source policy
passed for main and the inspected PR #134/#135 implementations.

Task identifiers can have an alphabetic suffix: `T-22.1a` is a distinct
originally completed task. Counting only numeric identifiers incorrectly
produces 82 original and 63 remaining entries.

Original acceptance criteria and definitions of done remain binding. The
reviewed implementation and its integration into `main` require separate
verdicts. Existing follow-ups for reopened entries remain owned by the
[tracking container](https://github.com/flyingrobots/keep/issues/132).

The [recovery findings in progress](recovery-findings.md) record a reproduced T-13.2 partial-seal contradiction that reaches a discard plan, owned by #171. This does not complete T-13.2's remaining criterion review or change the recorded verdict totals.
75 changes: 75 additions & 0 deletions docs/audits/completed-roadmap-criteria/architecture.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,75 @@
# Architecture decision and structural-enforcement audit

This page owns T-10.1 and T-10.2 from the originally checked roadmap at
`1a586d83d5750083172d440f90e7b786d540ff0e`, lines 496–513. Inspected main is
`f49cff732cf7a6e1b472decba9e4c4130990559e`. T-10.3's durable-port/fake inventory is evaluated below.

## T-10.1 — Decide the architecture

**Acceptance and delivery: met on main.** ADR-0004 is accepted and defines
inward dependency flow, semantic ports, codecs at boundaries, canonical
JSON/CBOR profiles and justified substitution boundaries. It records the
codec relocation, alternatives, invariants and compatibility consequences.
The task is the architecture decision; this verdict does not prove universal
implementation conformity or substitute for T-10.3's review.

## T-10.2 — Enforce structurally

**Acceptance/definition of done: not met.** The roadmap explicitly names
module size, forbidden filenames, no Python and denied unreachable public
visibility. Size and Python admission exist in `xtask/src/source_structure/`;
`Cargo.toml` denies `unreachable_pub`. The collector/classifier has no admission
for the nine Rust filenames prohibited by AGENTS.md.

A copy-isolated Docker main-equivalent clone with its own build directory
received one new inventoried `src/utils.rs` containing a module-ownership
comment. `cargo xtask source-structure-check` exited successfully. This is
observed acceptance of a prohibited name, not a compile error or zero-test
result. The temporary file and its index entry were removed afterward; the
clone returned to a clean state. No host worktree was mutated by the probe.

Correction owner: [issue #144](https://github.com/flyingrobots/keep/issues/144).
Its coherent scope is the existing literal basename prohibitions, with exact
refusals and preservation of current source/path/Python/size checks. It does
not introduce a new dependency-analysis requirement or substring naming ban.

## T-10.3 — Durable protocol ports and fault-injecting fakes

**Named acceptance and delivery: met on main.** The implemented durable write
protocols expose storage capabilities and deterministic fakes. The inventory
below checks the port and failure laws rather than relying on filenames alone.

| Protocol | Port | Executed failure evidence |
| --- | --- | --- |
| Immutable segment writing | `SegmentStage` | Scripted short/interrupted/zero/overreported writes and synchronization refusals; no receipt after failed durability |
| Store initialization | `StoreInitializationStorage` | Failure at each of six initialization phases; exact phase/source and attempted prefix |
| Catalog generation publication | `CatalogPublicationStorage` | Recording storage with exact phase failures and stopping later writes |
| Recovery stage discard | `RecoveryStageDiscardStorage` | Exact expected-state and directory-sync refusals; retained stage on refusal |
| Recovery stage completion | `RecoveryStageCompletionStorage` | Stage/pool/staging synchronization failures, pool conflict and operation-prefix assertions |
| Recovery next-head finalization | `RecoveryNextHeadFinalizationStorage` | Verification, candidate sync, replacement and root-sync failure laws |
| Recovery segment resume | `RecoverySegmentResumeStorage` | Injected storage failure returns no resumable stage; stale fingerprint refuses |
| Retention publication | `RetentionPublicationStorage` | All 17 publication phase failures preserve exact attempted prefix; authority failure precedes mutation |
| Store migration | `StoreMigrationStorage` | All 21 phase failures preserve exact attempted prefix; current-state verification failure precedes mutation |

The immutable segment port is exercised by the scripted `stage_double`;
catalog, recovery, retention and migration suites carry their recording or
in-memory storage doubles. These are substitution boundaries with observable
failures, not placeholder traits. Inspection of existing domain directories
found no adapter/filesystem/network/Serde imports, but this targeted review is
not an exhaustive architectural analysis of every boundary module.

Copy-isolated Docker with pinned Rust 1.96.0 and the dedicated main-equivalent
source build directory ran all eleven listed public integration targets: 72
laws passed in debug and 72 in release, with zero filtered or ignored tests.
The source clone has unchanged main Rust/corpus content; no new production
implementation or native test was introduced for this evidence.

This verdict establishes ports and fakes for implemented protocols. It does
not assert that passing mocks proves filesystem durability, that partial
migration recovery is delivered, or that retention production admission is
complete; those require their own later-task and correction-owner evidence.
T-10.2 remains unmerged on main, with its correction in PR #145.

Thirty remaining checked tasks now have verdicts. Thirty-four other checked
tasks and nineteen reopened entries still need full accounting. No checkbox
changed, and neither issue #131 nor its tracking parent is closed.
74 changes: 74 additions & 0 deletions docs/audits/completed-roadmap-criteria/benchmark-baseline.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# Benchmark-baseline acceptance audit

This page owns the originally completed T-09.1 verdict for issue #131.
Binding text is the original roadmap at
`1a586d83d5750083172d440f90e7b786d540ff0e`, lines 440–450.
Inspected main is `f49cff732cf7a6e1b472decba9e4c4130990559e`.
Originally unchecked T-09.2 and T-09.3 are excluded.

## T-09.1 — Corpus, scenarios, metrics and one baseline

**Named artifacts: present on main. Definition of done: not met.**

The roadmap delegates evidence to the benchmark README and supplies no
separate task-specific DoD block. The repository's parse/validate/admit and
canonical-format standards remain binding at the report publication boundary.
Correction owner: [issue #142](https://github.com/flyingrobots/keep/issues/142).

| Obligation | Inspected and executed evidence | Verdict |
| --- | --- | --- |
| Generated bounded corpus | `benchmark/src/corpus.rs`, generators and corpus tests: deterministic members, exact edit coordinates and identities, total byte limit 16,777,216 | Pass |
| Thirteen reproducible scenarios | Scenario catalog and scenario tests: frozen order, every scenario executes, deterministic semantic counters | Pass |
| Five profile comparisons | Profile tests: exact parameters and pinned provenance, exact partition coverage, edit reuse and production registered-profile identities | Pass |
| Required reference metrics | Measurement/report modules: wall/process-CPU time, nearest-rank percentiles, allocation/live-heap counters, bytes, operations and exact ratio fields | Pass for this reference scope |
| Verification mandatory; diagnostics distinguished from optimized evidence | Verification posture has no disabled state; build-profile law refuses debug publication; report metadata records posture | Pass |
| One source-bound committed baseline | `benchmark/baselines/c529c07-aarch64-apple-darwin.tsv`: existing Git commit c529c07, clean source, Rust 1.96.0, Darwin 25.3.0/M1 Pro, 100 samples/five warmups, 13 scenario and five profile rows | Artifact present; fields verified |
| Threshold policy explicit | Artifact and README mark all performance thresholds unconfigured; controlled-history work is originally unfinished T-09.2 | Pass |
| Bound captured source/compiler/host without ambiguity | Artifact validator accepts expected metadata plus a conflicting second git-commit coordinate | Fail |
| Admit complete canonical metric rows before publication | Existing accepted unit fixture contains bare scenario/index and profile/index rows, without headers or metric fields; validator checks counts only | Fail |
| Delivery | Existing implementation and historical artifact are on main; strict admission correction has not landed | Correction required |

The original artifact is a historical measurement witness, not proof that
current code has identical throughput. PR #134 has separately source-bound
single-pass evidence; cross-host timings cannot establish a speedup. This
audit neither repeats the historical timings nor invents hardware coordinates
for a new optimized baseline. Durable scenarios remain outside T-09.1.

## Reproduced failure

The production ingress is
`xtask/src/benchmark_baseline/artifact.rs::validate`. It checks that expected
metadata lines occur and that 13 scenario/five profile rows occur. It does
not exclude a conflicting metadata line or validate the complete row grammar.

In the accepted existing fixture, append:

```text
metadata<TAB>git-commit<TAB>ffffffffffffffffffffffffffffffffffffffff<LF>
```

The expected captured commit remains present as well. A Docker probe asserting
refusal failed at `conflicting source coordinates were admitted`: validation
returned success. No bad artifact was published. The probe was restored and
the source clone is clean. The first probe had a function-qualification compile
error and is not RED evidence; only the corrected running assertion is counted.

Issue #142 owns one coherent complete-report admission boundary, including
permanent typed mutation/fuzz evidence and protection of prior artifact state.
Existing positive tests and green builds do not discharge this missing law.

## Executed evidence and accounting

Copy-isolated Docker, pinned Rust 1.96.0, main-equivalent source clone
`7981988`, dedicated target directory:

- keep-benchmark: 19 unit laws and one public integration law passed in
debug and release.
- Two actual benchmark report-admission laws passed in debug and release.
- Conflicting-source probe: one executed law failed at the intended assertion.
- An earlier `benchmark_report` filter selected zero tests and is not evidence.

No native tests, fresh optimized baseline, full-workspace validation or
performance threshold claim is made. Twenty-seven checked tasks now have
verdicts; 37 other checked tasks and 19 reopened entries still need accounting.
T-06.3, T-06.4 and T-09.1 retain acceptance/DoD/delivery gaps. No checkbox changes.
Loading
Loading