Skip to content

test: verify catalog agreement with independent generated histories - #167

Merged
flyingrobots merged 5 commits into
mainfrom
test/166-generated-catalog-model
Oct 3, 2026
Merged

flyingrobots merged 5 commits into
mainfrom
test/166-generated-catalog-model

Conversation

@flyingrobots

@flyingrobots flyingrobots commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Landed

Merged as eb506dfb3830a32b0aec6a963c69da4f87012161 after fresh independent approval and all four hosted jobs passed for 239bd19553449f8e509cc7752f142f7e2c4e46ff. The signed integration commit preserves the approved tree. Final Code Lawyer closure supersedes the pre-landing status below.

Problem and outcome

The catalog model previously checked one fixed history and derived expected records through production segment decoding. It never queried absent identities: a deliberately broken lookup returning the first binding for an absent key passed that model. The generated model fails against the same defect and passes against unchanged production.

Closes #166. Refs #131. Current candidate 239bd19553449f8e509cc7752f142f7e2c4e46ff normally integrates main 5179ed78a74d19a3c24f300acbc5228144e6628a with no text conflicts.

Invariant and approach

Catalog snapshots return exactly the selected records and remain tied to their admitted generation. Enumerate bounded three-generation histories over input-derived chunk/layout maps with bundled and reversed separate segment packing. Compare exact payloads, absence, logical count and generation; independently compare successor coordinates and precise stale/skipped/predecessor/head refusals.

Change-Kind: test-evidence enhancement. Production code, format/API behavior and existing expectations are unchanged. Keep readable existing examples. Counting generated harness cases and deriving expectations through production catalog/segment iteration were rejected because neither supplies the intended independent runtime oracle.

Validation and review

Historical Docker debug/release model laws, Clippy, formatting, source policy and Markdown checks pass. Ten distinct production mutations produce runtime RED; the absent-lookup mutant also passes the old oracle. Compilation/setup failures and one cached-mutant run are excluded, with artifacts preserved. Prior independent review found a diagnostic-coordinate calibration gap; diagnostic-only controls then reached both exact coordinate assertions, with restored-production GREEN. Prior delta review approved c0e1a97d6ce4e57f493ac6da3fd54ee2d4ffaf2f, whose four hosted jobs passed in run 37090738323.

Those are historical results. Current integrated head 239bd19553449f8e509cc7752f142f7e2c4e46ff is receiving fresh full copied-Docker validation, exact-head independent review and hosted checks in run 37156173106. All current gates must pass before the authorized normal merge.

The evidence record specifies independent expectations, replay/reduction, mutation subjects and limitations. Exhaustiveness is limited to the declared finite input universe and history bound. It does not establish independent chunk hashing, arbitrary-length histories, concurrent scheduling, filesystem durability or newly enforced per-test resource ceilings.

Compatibility, recovery and security

No production behavior, public API, on-disk format, publication order, recovery protocol, dependency or security behavior changes. No benchmark impact is claimed. The in-memory serialization sink supplies test data and makes no physical durability promise. Existing infrastructure enforcement gaps remain disclosed, not waived. The mainline merge preserves previously reviewed sealed-stage, platform admission, recovery, reader-fence and migration evidence without changing their runtime implementations.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-03T02:38:36.332378Z 343a6f9 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Next included review available in 32 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: ASSERTIVE
  • Plan: Advanced
  • Run ID: 6ee86835-8794-4ee8-a846-fddc2894d47d
📥 Commits

Reviewing files that changed from the base of the PR and between 5179ed7 and 239bd19.

📒 Files selected for processing (6)
  • docs/formats/segment-store-v1/requirements.md
  • docs/testing-evidence/catalog-model-histories.md
  • tests/catalog_model.rs
  • tests/catalog_model/generated_histories.rs
  • tests/catalog_model/record_inputs.rs
  • tests/catalog_model/transition_refusals.rs
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@flyingrobots

Copy link
Copy Markdown
Owner Author

Independent review of PR #167

Reviewed exact pushed head 343a6f9ac91056bad67134f2dad377f67a5783da on test/166-generated-catalog-model against main 6051abb25a9fd33ae7ee0de5614514b709a4d82a in an isolated checkout. The checkout was clean and live GitHub head/base matched. This is the authorized independent Codex fallback under the adversarial agy protocol. No repository edits, host Rust execution, comments, commits, or subagents were used.

Finding

P2 — Calibrate the exact transition-refusal coordinate assertions

Mandatory evidence gap, not a demonstrated defect in production or the asserted expectations.

tests/catalog_model/transition_refusals.rs:44–51 promises the exact expected/observed generation coordinates, and :82–90 promises the exact predecessor digest coordinates. The supplied stale-generation and wrong-predecessor production mutants invert the refusal guards, causing unexpected success. Their red.log receipts stop in require_error with “stale/skipped generation admitted” and “wrong predecessor admitted,” respectively. Neither reaches the subsequent exact typed-error assertion.

These receipts correctly calibrate refusal existence, but not the distinct diagnostic-coordinate claims. A production regression that still refuses while swapping expected and observed fields is a different outcome; its protection is asserted in source but has no witnessed falsification here. Testing Standards rule 4 requires the named load-bearing assertion to execute and fail for the intended reason. The mixed-head receipt does reach its own typed-error assertion, but does not establish the two separate catalog-transition diagnostic oracles.

Narrow fix: add isolated diagnostic-only mutations preserving refusal while changing or swapping the generation and predecessor expected/observed coordinates. Record each existing exact assertion failing after compilation, then restore production and record relevant GREEN. Update the evidence table to distinguish refusal-existence calibration from diagnostic-coordinate calibration. No production change is requested unless an assertion actually survives.

Verification Checklist

Scope and implementation paths

  • Inspected all six changed files and the complete PR diff. Production and durable formats are unchanged. The existing handwritten model is preserved verbatim apart from loading the new test module.
  • Read applicable AGENTS, binding Testing Standards, enforcement-profile disclosures, issue Verify catalog model agreement over generated independent histories #166, PR text, requirement mapping, and the evidence record. The issue's observable outcome is generated independent catalog-model evidence; unrelated storage, concurrency, retention, and verification features are excluded.
  • Input oracle: tests/catalog_model/record_inputs.rs:18–45 defines two explicit chunk byte strings and two independently frozen layout identities/payload fixtures before segment encoding. The literal layout IDs agree with conformance/layout/v1/layouts.tsv:3–4. Expected membership never comes from production segment iteration or catalog enumeration. Shared chunk hashing is disclosed and does not establish independent hash correctness.
  • Membership generation: record_inputs.rs:47–59 uses checked mask-bit selection over the deterministic BTreeMap; generated_histories.rs:29–38 visits the fixed mask space. Zero masks produce empty catalogs; all subsets include additions, removals, unchanged membership, and changing layout membership within the declared universe.
  • Packing paths: record_inputs.rs:68–80 emits bundled records or reversed separately packed records. :82–103 uses the actual segment staging/record constructors; :106–118 is an owned in-memory serialization sink. It makes no synchronization or persistence claim. Layout re-encoding feeds production bytes, while expectations remain the original literal fixture bytes.
  • Construction/admission: generated_histories.rs:40–72 creates models before encoding, uses AdmittedSegment::decode, constructs catalogs/heads, and admits snapshots through public APIs. Traced the relevant production boundaries in catalog_encoder.rs:14, catalog_admission.rs:10, and catalog_snapshot_admission.rs:5. Test expectations are not reconstructed from those outputs.
  • Pinned lookup/count/generation: generated_histories.rs:73–83 checks all retained snapshots after their construction; :128–151 checks count and exact payload or absence for every input identity. Production runs through catalog_snapshot.rs:20/41/47 and admitted_catalog.rs:42/49. These are in-memory pinned-view checks, not mutable-filesystem publication or concurrent-reader proofs.
  • Valid successor: generated_histories.rs:101–125 compares public successor results with independently derived generation arithmetic, tracing admitted_catalog.rs:69 to catalog_transition.rs:8/22 and catalog_successor.rs:17.
  • Refusal paths: transition_refusals.rs:16–54 exercises stale/skipped generations; :58–91 exercises wrong predecessor digest; :95–125 exercises head/catalog generation mismatch. The first two trace the exact transition guards; the third traces catalog_snapshot_admission.rs:5. Expected typed errors are correct by inspection. Their calibration status is distinguished in the finding.
  • Existing example: tests/catalog_model.rs:26–50 still checks its original three-generation example. Its model remains correlated with segment decoding, but that limitation no longer describes the new oracle. Retaining it does not substitute for the generated evidence.

Generation, reduction, constants, and claims

  • Four fixed records imply 16 masks. Three nested mask loops produce 4,096 histories per packing, with two packing modes. These are bounded enumeration facts, not evidence of correctness by case count. The record universe is fixed, not arbitrary payloads/layouts or unbounded histories.
  • Generation expectations are 1/2/3 for successful histories; stale/skipped candidates are 2/4 against expected 3; the mixed-head law expects head generation 1 versus catalog generation 2. Checked conversions/arithmetic are used where derived values cross representation boundaries.
  • The literal layout record lengths 176 and 220 and their identities match frozen corpus rows. The declared policy maxima are existing production bounds, not newly claimed measurements or stress coverage.
  • Exhaustive replay follows deterministic packing/mask order and stops at the first failure. It supplies an explicit bounded counterexample order; it is not a general arbitrary-length shrinker. No random seed is needed for this fixed enumeration. Raw replay files identify source commit and exact test command outside the child process. New production counterexamples would still require permanent named regressions; none is claimed here.
  • The model's enumeration covers changes of layout membership. It does not claim multiple admissible representations of the same blob, retained closure, natural chunk boundaries, arbitrary scheduler exploration, or crash durability.
  • The proposed small classification is based on memory-only resource topology. Existing runner isolation, latency-budget, and resource-ceiling gaps remain explicit; neither this review nor the gap ledger waives them or claims enforcement has been implemented.

Calibration and execution evidence

  • Inspected all eight production mutation definitions, their replay coordinates, and their compiled runtime failures under scratch issue 166. Missing lookup, unexpected lookup, count, pinned generation, and successor generation each reach the corresponding exact observation assertion. Stale-generation and wrong-predecessor mutations reach refusal-existence checks only. Mixed-head reaches its exact typed-error assertion, observing CatalogDigest where Generation is required.
  • unexpected-record/red.log records Some([0]) versus None at history [0, 0, 1]; missing-record/red.log records the inverse presence failure. old-oracle-survives.log records the original named example passing against the absent-lookup mutant. This is a calibrated evidence enhancement, not a claim that main contained the mutant.
  • record-count, pinned-generation, and successor-generation receipts show the intended zero/one substitutions failing against input-derived expectations. No timing or mutation-percentage gate is claimed.
  • The missing-record-compile-failure receipt is a dead-code compilation failure and is excluded. The corrected mutation retains the record call and filters its result, reaching the runtime payload assertion.
  • final-focused-green.log contains the disclosed cached-mutant failure and is not accepted as candidate GREEN. final-focused-clean-green.log records package artifact removal, recompilation, and five passing tests in debug/release; the subsequent source-policy command fails because the copied tree lacks Git metadata. That setup failure is not product evidence.
  • verified-green.log records the unmutated five-test suite passing debug/release after the documented correction. The original generated-suite receipts under scratch issue 131 remain historical execution evidence. No unrelated setup failure is counted as RED, and observed durations are not promoted into latency promises.
  • Compared calibrated test commit c9b41e99fb7e89217afbd896732dbea09a0c97ab with the final test files: the only subsequent test change is formatting the successor-expression assertion. Its behavior and expected values are unchanged. Documentation changes do not silently expand calibrated runtime scope.

History, review surfaces, and pending checks

  • Inspected the three PR commits: test addition, evidence documentation, and Markdown whitespace correction. There are no merge commits in the PR; no merge-parent integration resolution is unreviewed.
  • Live review-thread and review connections are empty, with hasNextPage: false. Read both global comments completely; the global-comment connection also reports no further page. Hosted Codex review is running; CodeRabbit reports a review limit. Its success status is not treated as substantive approval.
  • Live required checks for exact head 343a6f9 show documentation/workflow integrity and dependency policy passed; Rust quality gates and runtime fuzz smoke are still in progress. Pending checks are a separate readiness condition, not the reason for the calibration finding.
  • The changed requirement row links generated evidence without claiming new production behavior. The eight-mutation table describes actual runtime REDs, but needs the additional direct diagnostic calibrations before the complete assertion-calibration requirement is satisfied.

Execution limits and verdict

This reviewer executed only read-only source/history/log inspections and live GitHub queries. Tests and mutations were inspected from raw Docker receipts, not rerun by the reviewer. Static review and finite generated evidence do not prove absence of all regressions, arbitrary-input behavior, independent hashing, scheduler correctness, or physical durability.

No demonstrated production defect was found. Approval is withheld for the narrow mandatory evidence gap above. The verdict applies only to the exact reviewed head; approval would not itself authorize merging.

REQUEST CHANGES — 343a6f9ac91056bad67134f2dad377f67a5783da.

@flyingrobots

Copy link
Copy Markdown
Owner Author

Independent final delta review of PR #167

Reviewed exact pushed head c0e1a97d6ce4e57f493ac6da3fd54ee2d4ffaf2f, remediation parent 343a6f9ac91056bad67134f2dad377f67a5783da, against main 6051abb25a9fd33ae7ee0de5614514b709a4d82a in an isolated checkout. Local and live GitHub coordinates match, and the checkout is clean. This is the authorized independent Codex fallback under the adversarial agy protocol.

Findings and disposition

No remaining actionable finding in the reviewed candidate.

The prior P2 calibration gap is closed. The generation-coordinates mutation preserves refusal and swaps generation coordinates in production catalog_transition.rs. Its runtime log reaches tests/catalog_model/transition_refusals.rs:44, failing with expected/observed 2/3 instead of the specified 3/2 at mask 0, candidate 2.

The predecessor-coordinates mutation likewise preserves refusal and swaps the predecessor coordinates. Its runtime log reaches transition_refusals.rs:82, failing the exact typed predecessor assertion at mask 0. Both runs compile successfully and execute the named test; neither stops at the earlier unexpected-success check.

diagnostic-restored-green.log records recompilation after restoring production, followed by all five catalog-model tests passing in debug and release. The only tracked delta adds the two calibration rows and their explanation to the evidence document. Production and tests are byte-identical to the previously reviewed head; no expectation or storage behavior was changed.

Verification Checklist

Scope, paths, and integration

  • Inspected the complete delta and verified that only docs/testing-evidence/catalog-model-histories.md changed. The previous full-PR source review remains applicable; the prior finding is resolved by new direct evidence, not a weaker test or reduced contract.
  • Input oracle retained: record_inputs.rs:18–59 constructs chunk/layout expectations and mask-selected membership before encoding. Frozen layout identities/payloads remain independently specified; shared production chunk hashing remains a disclosed limitation.
  • Packing/construction retained: record_inputs.rs:68–118 exercises bundled and reversed separate packing through real segment staging with an in-memory sink. generated_histories.rs:40–72 uses public segment, catalog, head, and snapshot admission. Expectations remain separate from returned catalog records.
  • Membership and pinning retained: generated_histories.rs:73–83/128–151 checks generation, exact payload, absence, and count after snapshots are constructed. Production paths remain catalog_snapshot.rs:20/41/47 and admitted_catalog.rs:42/49.
  • Successor retained: generated_histories.rs:101–125 compares independently derived successor generations through admitted_catalog.rs:69, catalog_transition.rs:8/22, and catalog_successor.rs:17.
  • Generation refusal calibrated: transition_refusals.rs:16–54 reaches catalog_transition.rs:29–30; the new diagnostic mutant leaves the mismatch guard and refusal intact, then swaps only error fields. The exact assertion at line 44 now has direct runtime falsification in addition to earlier refusal-existence evidence.
  • Predecessor refusal calibrated: transition_refusals.rs:58–91 reaches catalog_transition.rs:34–35; the new mutant preserves refusal and changes only diagnostic coordinates. The exact assertion at line 82 now has direct runtime falsification.
  • Head binding retained: transition_refusals.rs:95–125 and catalog_snapshot_admission.rs:5 retain the previously calibrated generation-mismatch oracle. No new or parallel path bypasses the existing checks.
  • State-machine/architecture continuity: no production, public API, codec, format, resource limit, writer, recovery, cancellation, or durability path changed. The old handwritten model and all original expected values remain intact.
  • No merge commits occur in the remediation delta. The prior three-commit source/history review and its no-merge finding remain applicable.

Evidence and numerical claims

  • Inspected both new original.rs/mutant.rs diffs, replay.txt files, and compiled runtime red.log failures. Replay records pin test head 343a6f9, source path, and exact test command. Because this final delta changes documentation only, those tests are identical at the reviewed resulting head.
  • Inspected diagnostic-restored-green.log: compilation precedes the debug run, all five tests pass, then the release build and all five tests pass. These are inspected Docker executions, not reviewer-run tests.
  • The evidence table now separates refusal-existence calibration from expected/observed-coordinate calibration. Its statements agree with the actual failing assertion locations and outcomes.
  • Earlier eight mutation receipts remain binding: missing/present lookup, record count, pinned generation, successor generation, stale/skipped refusal, predecessor refusal, and mixed-head diagnostics. The old model's survival against absent-lookup substitution and the generated model's failure remain valid evidence of the intended improvement.
  • No new numeric limit or performance claim is introduced. The four-record, 16-mask, three-generation, two-packing bounds and frozen layout lengths/identities are unchanged. Exhaustive replay remains bounded by that input universe and deterministic order; it is not an arbitrary-history reducer or universal proof.
  • Previously excluded compile failure, stale-artifact run, and missing-Git-metadata setup failure remain excluded. No setup result is reclassified as runtime calibration. Existing resource-ceiling and isolation enforcement gaps remain disclosed rather than waived.
  • Applicable AGENTS and Testing Standards still govern the work. The two previously uncalibrated load-bearing assertions now each have their own intended runtime RED and restored GREEN. No new test assertions were introduced by this delta.

Review surfaces and execution status

  • The prior full review was posted before remediation at PR comment 5964746505. Rechecked live review threads, reviews, and global comments; every queried connection reports hasNextPage: false. No new actionable review finding is present. Hosted Codex shows completion for the earlier head; CodeRabbit remains rate-limited. Neither status is substituted for this exact-head review.
  • Live GitHub confirms head c0e1a97d6ce4e57f493ac6da3fd54ee2d4ffaf2f and base 6051abb25a9fd33ae7ee0de5614514b709a4d82a. At inspection, dependency policy passed for run 37090738323; Rust quality gates, documentation/workflow integrity, and runtime fuzz smoke were still in progress. All required checks must complete before readiness; no prior-head green is transferred.
  • Executed read-only Git/source/log inspection and GitHub queries only. No host Rust, new mutation execution, repository changes, external comments, or subagents were used. Runtime evidence is explicitly inspected-only.

Limits and verdict

The finite model does not establish independent chunk hashing, arbitrary record/history spaces, concurrent scheduling, filesystem publication, or physical durability. The replay order and in-memory sink retain their documented scope. Static review and green finite tests do not prove absence of all defects.

The sole prior finding is closed. This approval applies only to the exact resulting head, leaves required current-head CI as a readiness gate, and does not authorize merging.

APPROVE — c0e1a97d6ce4e57f493ac6da3fd54ee2d4ffaf2f.

@flyingrobots

Copy link
Copy Markdown
Owner Author

Code Lawyer activity summary

Candidate: c0e1a97d6ce4e57f493ac6da3fd54ee2d4ffaf2f, based directly on main 6051abb25a9fd33ae7ee0de5614514b709a4d82a. Production files and durable formats are unchanged.

Obligation Evidence and disposition
Independent generated model agreement (#166) Input maps precede production encoding; exact payload, absence, logical count, pinned generation, successful successor and typed refusal outcomes. Bounded histories and both packing paths are documented; no harness-count assertion.
Load-bearing runtime calibration Production mutations break lookup presence/absence, count, snapshot generation, successor generation, refusal existence and head binding. Each has a compiled runtime RED. The original model survives the absent-record mutant; the new one fails.
P2 independent-review finding: diagnostic coordinates uncalibrated Full initial review posted before remediation. Diagnostic-only generation/predecessor mutations preserve refusal and fail the exact assertions at transition_refusals.rs:44 and :82. Restored production passes debug/release. Commit c0e1a97 records closure; no implementation change was needed.
Setup failures and cache integrity Dead-code mutant compilation, missing copied Git metadata, and a cached-mutant post-restore run are excluded. Corrected receipts retain original failures separately and establish successful rebuilt production execution.
Limits Fixed input universe and bounded histories; shared chunk hashing is disclosed. No scheduler, filesystem durability, arbitrary-length history, or new resource-ceiling enforcement claim. These limits are not risk waivers.

Local Docker verification: model laws in debug/release, focused Clippy, formatting, source policy and Markdown lint pass. No mutable host checkout was mounted into the test runner.

Final independent review and final-head hosted checks are recorded below. CodeRabbit's rate-limit success status is not substantive review approval. This PR is not merged; review approval does not itself authorize a merge.

Final gate: independent Codex reviewer /root/keep164_independent_review APPROVE with mandatory checklist for c0e1a97d6ce4e57f493ac6da3fd54ee2d4ffaf2f. CI run 37090738323 passes Rust quality, documentation/workflow integrity, runtime fuzz smoke and dependency policy on that same SHA. The final paginated review-queue read reports no inline threads, no review objects and no additional comment pages; the prior independent finding is explicitly superseded by the final review. Local and pushed heads match and the worktree is clean.

READY FOR MERGE, subject to maintainer merge authorization and repository protections. The PR is already out of draft. No merge was performed.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@flyingrobots

Copy link
Copy Markdown
Owner Author

Independent exact-head review of PR #167

Reviewed flyingrobots/keep, branch test/166-generated-catalog-model, head 239bd19553449f8e509cc7752f142f7e2c4e46ff, tree ee612cb20dc1e7db0299411fb570e2275a28c1e1, target main 5179ed78a74d19a3c24f300acbc5228144e6628a. The isolated checkout is clean. This is the explicitly authorized independent Codex fallback under the ULTRA STRICT agy protocol. The review is read-only except for this scratch report; no source edits, host Rust execution, comments, commits, configuration changes or subagents occurred.

Findings

No verified actionable P0–P5 issue in the reviewed change. The previous diagnostic-calibration finding remains closed by direct runtime evidence. No production defect is inferred from calibration mutants.

This is an exact-head source/integration approval. The parent agent reports terminal EXIT 0 for its full Docker chain on this exact tree, and this reviewer inspected the completed log including debug/release model outcomes and final fuzz Clippy completion. Current-head hosted checks and repository protections remain separate readiness gates. No hosted check outcome for this head is claimed here.

Verification Checklist

Scope and prior work

  • Inspected AGENTS.md, binding Testing Standards, enforcement profile and relevant documentation standards. The complete six-file diff against the actual target changes the catalog requirement's evidence routing, adds the finite model and its evidence, and loads it alongside the existing example. No production, dependency, format, API, fuzz implementation or crash implementation differs from target main.
  • Verified tests/catalog_model.rs, all three new test files and docs/testing-evidence/catalog-model-histories.md are byte-identical to approved c0e1a97d6ce4e57f493ac6da3fd54ee2d4ffaf2f. Read the earlier independent review and independently rechecked the underlying source and raw calibration, rather than inheriting its verdict by assertion. Existing example expectations are preserved; the only original test-file addition is module loading.
  • Change declaration is an evidence enhancement with no changed product promise or expectation. Subject is actual public catalog behavior; mutation receipts are calibration, the in-memory sink is a serialization fixture, and neither harness counts nor tool success replace product outcomes. There is no claimed product bug fix requiring fabricated RED on unmodified main. The evidence records owner, oracle limits, deletion criteria, fixed replay order, source pins, debug/release profile and missing resource enforcement. Existing per-test budget/sandbox enforcement gaps remain explicit gaps, not waivers or claims of compliance.

Every changed behavioral path and relevant parallel path

  • Input expectations: tests/catalog_model/record_inputs.rs:18–44 establishes two explicit chunk byte strings and two literal layout records/IDs before encoding. The frozen IDs agree with conformance/layout/v1/layouts.tsv:3–4; literal records have 176 and 220 bytes. The shared chunk-hash implementation is explicitly disclosed; independent hash correctness is not a claim of this model.
  • Membership selection: record_inputs.rs:46–58 selects checked mask bits over deterministic BTreeMap order; generated_histories.rs:29–37 enumerates every tuple of three masks. Empty, added, removed, unchanged and layout-containing states are represented. No expected membership is obtained from catalog or segment iteration.
  • Both packing paths: record_inputs.rs:65–76 emits bundled records, zero segments for empty membership, or reversed separately packed records. :78–99 uses real StagedSegment::begin/append/seal with admitted chunk/layout constructors. :102–118 supplies an owned in-memory byte sink with no claimed durability. Production staging is src/adapters/staged_segment.rs:43/80/113; sealing consumes staged state and propagates exact phase errors. Layout decode/re-encode supplies actual bytes while the expected map retains the independent literal fixture bytes. All generated inputs are bounded and valid; no recovery/fault injection is claimed by the sink.
  • Construction and admission: generated_histories.rs:40–72/86–99 builds input maps first, then encoded segments, catalogs, heads and all snapshots through public admission. Production runs through src/adapters/admitted_segment.rs:32, catalog_encoder.rs:14–32/50–79/127–138, catalog_admission.rs:10–42/68–105/108–124, and catalog_snapshot_admission.rs:5–24. Sorting, duplicate refusal, exact segment/record identity/checksum binding and head generation/length/digest admission remain present. Checked generation arithmetic does not admit overflow.
  • Pinned generation/count/lookup: generated_histories.rs:73–83/132–154 queries all retained snapshots after they have been created, checking each generation, logical count and every input identity's exact payload or absence. Production routes through catalog_snapshot.rs:21/42/48 to admitted_catalog.rs:43/49–60. Failed binary search returns absence; no fallback record is substituted. Count plus complete known-identity membership protects against omitted/extra records within the declared input universe. This checks immutable borrowed views, not a changing filesystem head.
  • Successful successor: generated_histories.rs:102–129 exercises both adjacent transitions with expected generations derived from input index arithmetic. Production is admitted_catalog.rs:70–74 → catalog_transition.rs:6–16/19–37 → catalog_successor.rs:15–16. Candidate generation and predecessor must be exact before the candidate proof is returned.
  • Stale/skipped refusal: transition_refusals.rs:16–54 uses every mask with candidates 2 and 4 after current generation 2, requiring generation 3 in the exact typed error. Production is catalog_transition.rs:25–30. Both refusal existence and the existing coordinate equality assertion are directly calibrated.
  • Wrong predecessor refusal: transition_refusals.rs:58–91 uses every mask with a generation-3 candidate naming the generation-1 digest rather than the current generation-2 digest. Production is catalog_transition.rs:32–35. Refusal and precise expected/observed digest fields are independently calibrated.
  • Mixed head/catalog generation: transition_refusals.rs:95–125 binds a generation-1 head to a generation-2 catalog and requires the exact 1/2 error. Production is catalog_snapshot_admission.rs:9–12, ahead of length/digest checks. The mutation receipt reaches the typed assertion and observes the wrong later digest refusal.
  • Existing parallel example: tests/catalog_model.rs:26–49/52–102 retains the original decoded-segment-derived model and named example. Its oracle correlation is preserved as an existing example, not promoted to independence. The absent-lookup mutation demonstrates the new model's additional protection.
  • Filesystem parallel admission: filesystem_catalog_snapshot.rs:61–96 reconstructs admitted segments/catalog/head from retained bytes and invokes the same catalog/head admission; restart coordinate checks are catalog_restart_loader.rs:94–112. catalog_publication_execution.rs:11–27/30–122 retains verified-current and ordered publication paths. The new tests do not exercise filesystem publication or assert path parity beyond the shared admission boundary; their documentation does not claim otherwise.
  • Errors, resource bounds and state: real production refusals propagate through ?; test boxing preserves sources. Required failure assertions use require_error then exact typed equality. No caller can continue a consumed failed stage. No async cancellation, shutdown hold, configuration switch, new lock or durable transition is introduced. Crash/recovery/concurrency/power-loss coverage remains owned by the preexisting suites, rather than attributed to this model.

Both-parent merges and incoming contracts

The new merge 239bd19553449f8e509cc7752f142f7e2c4e46ff has parents c0e1a97d6ce4e57f493ac6da3fd54ee2d4ffaf2f and 5179ed78a74d19a3c24f300acbc5228144e6628a. Inspected both-parent path diffs and the combined diff. Relative to its first parent it imports 136 mainline paths; relative to its second parent it contains exactly the six topic paths. Every incoming production/test/receipt path is therefore exactly the target-main blob. The combined requirement ledger retains both the strengthened platform-admission row and generated-model row; a clean text merge was not used as a semantic proof.

Inspected both-parent path inventories and combined integration hunks for the following incoming merges, with the contract checks below. Short SHAs uniquely identify full commits in the inspected graph; these are preserved incoming integrations, not new topic-owned production changes.

Merge Integration obligation inspected
d0cff10 (#172) Partial-seal fixed-framing refusals and the emitted runtime counterexample remain.
e781c0b, 182e495 (#158) Sealed mutable escape is removed; observation owns stages privately; crash callers remain migrated.
657593f, 64fafe3 (#157) Legacy repository publisher route retains actual production profile admission.
0565879, 1325841 (#156) Public-stage evidence and distinct phase-injection responsibility are retained.
64bbbf9, d08fafb (#159) Operation-derived release/restore model and exact refusal coverage remain.
ee21b01, 34d7090 (#160) Reader process evidence retains real kernel ordering and its scope.
1c2b9d7 (#175) Reader fixture retains migration writer authority continuously.
ff6f5be, 2f22d0d, 80d23f5 (#161) Complete migration namespace verification and lawful retained content remain; reader fixture repair is retained.
6e7e2d6, 5179ed7 (#162) Compatibility/parser/planner evidence retains the stronger emitted partial-seal witness and #161 completion correction.

Verified relevant present source, rather than trusting “adopts” text: sealed_segment.rs:8–22/69–77 exposes no arbitrary public mutable conversion; the new sink never calls one. filesystem_catalog_publisher.rs:94–101 delegates through FilesystemPlatformAdmission::from_repository_writer_lock, preserving profile checks; catalog modeling does not acquire publisher authority. reader_fence_process/fixture.rs:32–41 returns the retained migration authority. migration_recovery_execution.rs:145–164 invokes verify_complete before reporting Complete; filesystem_migration_recovery.rs:86–88 supplies version-two namespace admission. xtask/src/fuzz_seed_corpus/tests/materialization.rs:49–54 runs the emitted counterexample before deterministic rematerialization. These paths and all corresponding incoming receipts are byte-identical to target main.

The #99 ledger's incomplete-stage preservation/explicit-disposition deferral, cooperating-writer scope and precise known/uncertain failure-effects contract remain unchanged (docs/testing-evidence/retention-landing.md:9–13). No catalog model assertion expands or weakens these contracts. Incoming historical subsystem measurements are preserved with their owning receipts; this bounded catalog review does not newly certify every historical numerical claim in unrelated mainline documents.

Every new constant and numeric/document claim

  • Four distinct records, 16 membership masks, three generations, two packing modes agree with source. This gives 4,096 histories per packing and 8,192 executions of check_history by derivation, not by a case-count correctness assertion. Negative laws cover all 16 masks; stale/skipped candidates are 2 and 4 with expected 3. Generation/index additions and mask shifts use checked operations/TryFrom; windows have width 2. No timing, rate, buffer, timeout or memory budget constant is introduced.
  • MAXIMUM record/layout policies are unchanged protocol bounds (1,048,576 each), used only for these tiny inputs. No new performance threshold or allocation claim is made. Record fixture lengths/IDs, Rust 1.96.0, source SHA c9b41e99..., diagnostic source SHA 343a6f9..., issue/requirement routing and all mutation-table rows were checked against actual source/fixtures/raw receipts.
  • The evidence's prose paragraphs each occupy one physical line. It documents a finite lexicographic replay reducer and explicitly excludes arbitrary-length/minimal semantic reduction; no random seed or accumulating-seed lane is applicable to this exhaustive fixed universe. No changed golden vector or expectation baseline, deletion, or benchmark exists. New files are 158, 118 and 126 lines; largest new function is below the 60-line hard limit, parameter count is at most five, and nesting is at most three.
  • Raw receipt inventory has ten accepted runtime controls total: eight original controls plus two diagnostic controls, and one excluded compilation-only attempt. The PR body and evidence table's ten-control count are consistent. “Ten plus two” would double-count the diagnostics.

Calibration and raw evidence coordinates

Inspected each control's original.rs, mutant.rs, replay.txt, exit.txt and runtime failure in author-evidence/keep-audit/166/<control>/. Every accepted control exited 101 after compilation and running its named test. Compared every original.rs byte-for-byte with current production: all match. Test files are unchanged from the recorded source pins; historical calibration therefore applies to the identical assertions and touched production paths at this head.

Control Witnessed failure
missing-record Exact payload/absence, history [0,0,1], generation 3: None instead of [0].
unexpected-record Exact absence, same history: [0] for the absent second chunk.
record-count Membership count 0 instead of 1.
pinned-generation Pinned generation 1 instead of 2.
successor-generation Successor generation 1 instead of 2.
stale-generation Refused-operation check observes unexpected success; does not alone calibrate fields.
wrong-predecessor Refused-operation check observes unexpected success; does not alone calibrate fields.
generation-coordinates Exact assertion transition_refusals.rs:44 observes swapped 2/3 instead of 3/2.
predecessor-coordinates Exact assertion transition_refusals.rs:82 observes swapped digest coordinates.
mixed-head Exact assertion transition_refusals.rs:116 observes CatalogDigest instead of Generation 1/2.

old-oracle-survives.log compiles/runs the original model and passes against the absent-lookup mutant. diagnostic-restored-green.log recompiles restored production and records all five tests passing in debug/release. verified-green.log records Clippy and rebuilt five-test debug/release passes. final-focused-clean-green.log additionally retains the missing-Git-metadata source-policy setup failure; it is not represented as a completed full validation. The compile-only missing-record attempt and cached-mutant restoration attempt remain excluded as documented, not reclassified as calibration.

Review surfaces, execution and limits

  • Inspected the complete supplied 167-queue.json: five global comments, zero reviews and zero inline threads. The helper's pagination loops cover every connection, including thread comments. Earlier finding and final disposition are explicit; CodeRabbit's rate-limit status and hosted Codex's old-head completion are not this review's approval. No outstanding actionable thread is hidden by treating global comments as resolvable threads.
  • Executed read-only Git/source/receipt inspection, blob comparisons, independent fixture-length/count derivation and git diff --check only. No new Rust test, mutation or host build was executed by this reviewer.
  • Inspected the completed parent-run 167-validation.log: startup pins copied source tree ee612cb20dc1e7db0299411fb570e2275a28c1e1, Rust 1.96.0, Linux aarch64 and ext4 scratch. Commands include worldline and conformance checks, production durability crash campaigns in debug/release, source structure, formatting, workspace feature checks/Clippy, workspace debug/release tests, doctests, rustdoc and fuzz formatting/check/Clippy. All five catalog-model tests pass in both profiles (:556–565 and :2617–2626); the final fuzz Clippy completes successfully. The parent supplies terminal session 64980 EXIT 0. This is inspected execution plus the runner owner's terminal-status attestation, not execution by this reviewer. Dependency policy and runtime fuzz smoke are not claimed from this local log; final hosted checks and merge protection/authorization checks must still be established. Historical focused green is not substituted for current-head full green.
  • Scope gaps are explicit: independent chunk hashing, arbitrary record/history spaces, arbitrary absent identities outside the universe, concurrent filesystem head changes, scheduler exploration, physical durability/power loss, and automated ordinary-test resource ceilings are not established by this model. Unrelated historical subsystem receipts were preserved and relevant integration contracts checked, not rerun or universally recertified. Static inspection and finite tests do not prove absence of all defects.

APPROVE — 239bd19553449f8e509cc7752f142f7e2c4e46ff

@flyingrobots

Copy link
Copy Markdown
Owner Author

Code Lawyer activity summary — ready to merge

Candidate 239bd19553449f8e509cc7752f142f7e2c4e46ff, tree ee612cb20dc1e7db0299411fb570e2275a28c1e1, integrates main 5179ed78a74d19a3c24f300acbc5228144e6628a with both-parent integration inspected. Change kind: test/evidence enhancement; production behavior, API and format are unchanged.

Item Severity / source Disposition and evidence
Catalog membership and transition evidence #166 Public catalog observations are checked against input-derived finite histories, with exact payload, absence, count, generation and typed transition outcomes. Existing examples remain.
Diagnostic calibration Prior independent P2 Closed: ten accepted runtime controls total, including the two exact diagnostic-coordinate controls. Raw originals and test blobs still match current source. Compilation/cache/setup failures remain excluded from evidence.
Mainline integration Landing review Incoming runtime paths and evidence remain intact, including #161 completion admission, #162 emitted counterexample, and #175 continuous writer authority.
Independent acceptance Authorized Codex fallback Fresh read-only GPT-6.1 high-reasoning APPROVE and full checklist cover this exact candidate.
Local validation Current candidate Full copied-source Docker chain passed on the exact tree: worldline/conformance/structure, both crash campaigns, feature checks, formatting/Clippy, debug/release suites, doctests/rustdoc, pinned toolchain, and fuzz build/Clippy.
Hosted validation Current candidate All four jobs in run 37156173106 passed. CodeRabbit is rate-limited, not an approval.
Complete review queue Current landing Earlier fully paginated GraphQL snapshot had no reviews or inline threads. GraphQL refresh returned 504; fully paginated REST refresh inspected all six conversation comments, zero reviews and zero inline comments. Only the hosted Codex quota notice was new; no actionable finding remains.
PR description Workflow correction A failed GraphQL read briefly led to clearing the body. Full body was restored and verified through REST; no source was changed.

This finite model does not establish independent chunk hashing, arbitrary input/history spaces, concurrency, filesystem publication, power-loss durability or automatically enforced per-test resource ceilings. Historical receipts retain their original coordinates and exclusions. The tests account for runtime claims; history counts describe the explored space and are not correctness assertions.

MERGE GATE: OPEN. Maintainer authorization already covers normal merging after clean review and green validation. Active repository signature/history protections remain enforced.

@flyingrobots
flyingrobots merged commit eb506df into main Oct 3, 2026
5 checks passed
@flyingrobots
flyingrobots deleted the test/166-generated-catalog-model branch October 3, 2026 21:55
flyingrobots added a commit that referenced this pull request Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Verify catalog model agreement over generated independent histories

1 participant