Conversation
…orphan path
Dashboard DELETEs create kind=erase obligations under an erasure id
with no ingestion_operations row. The 30s redrive fed them to the
ingestion-only context lookup and failed them as orphans
('canonical record missing') until exhausted, so successful deletes
surfaced as DEGRADED. The redrive now proofs erase groups read-only
and marks ERASED only on full proof, never bumps attempts, and the
orphan path only touches apply obligations.
Collaborator
|
Merged into 4.1.21, now on PyPI and npm, with your authorship kept. Thank you, this was a real bug. On top of your fix, unconfirmed deletions now back off instead of retrying forever, they are re-checked in every profile, entity erasures heal, and an unreadable vector store never counts as erased. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
A successful dashboard DELETE surfaced as
DEGRADEDwith an unfixable failure. Deleting factff8a4c44created threekind=eraseobligations under erasure id210a524c(one per owner:bm25/temporal/vector). The 30s background_reconcile_pending_projectionsfeeds every pending operation id to the ingestion-only_context_for_operation, which looks upingestion_operations— a table erasure ids never appear in. Context resolves toNone, so_terminalize_orphan_operationFAILED all three obligations with{"phase":"orphan","error":"canonical record missing"}and bumped attempts. After 10 passes the op was exhausted.retrycan never succeed: there is no ingestion row to reconcile. Error text is misleading twice over — the missing row is the ingestion record, and the fact itself was correctly deleted and tombstoned.Solution
Fan the redrive out by obligation kind instead of treating every operation as an ingestion:
_pending_obligation_kindsreads the DISTINCT non-terminal kinds per op (loud on query error — never a silent skip).erase-kind groups go to a new_reconcile_erase_operation: requires canonical absent (profile-scoped) + tombstone present + every pending owner proving no residue via read-onlyErasureService.prove_erased(new public wrapper over_prove_owner). Full proof →ERASEDwith no attempt bump, plus a manifest write so completed erasures leave the missing-manifest feed. Anything unproven stayspending, untouched._terminalize_orphan_operationnow only toucheskind=applyand saysingestion record missing.Alternatives rejected: bumping attempts on unproven erasures (ages them into silent abandonment); marking
ERASEDon tombstone alone (tombstone is written before owner purge inremove(), so presence proves nothing about residue); filteringoperations_missing_manifestSQL (shared surface, wider blast radius — handled instead by skipping empty-kind ops and writing manifests on close).Call tree:
Changes
src/superlocalmemory/server/unified_daemon.py— the fix_pending_obligation_kinds()_reconcile_erase_operation()+_reconcile_erase_group()ERASED, no bump, manifest on close_terminalize_orphan_operationskips non-apply, renames errorsrc/superlocalmemory/core/transactions/erasure.py— additive onlyErasureService.prove_erased()public wrapper_prove_ownerinternalstests/core/test_projection_spine_integration.py— regression teststest_redrive_does_not_orphan_erase_obligations— erase-only op + tombstone, no ingestion row → allerased, attempts 0, manifest writtentest_redrive_leaves_erase_pending_when_residue_remains— bm25 row present → stayspending, attempts 0test_redrive_leaves_erase_pending_without_tombstone— no tombstone → stayspending, attempts 0 across 2 passesWhat Does Not Change
finalizealready emits the receipt path); no receipt rows written by the redrive.pendingwithout bumping — visible but neverDEGRADED; follow-up may add observability/aging.Breaking Changes
None.
Test Plan
Regression test (fails before fix, passes after):
tests/core/test_projection_spine_integration.py::test_redrive_does_not_orphan_erase_obligationsAssertionError: bm25: {'owner': 'bm25', 'state': 'failed', 'attempts': 1} — assert 'failed' == 'erased'erased, attempts 0, manifest present)Existing suites (no regression):
pytest tests/core/test_projection_spine_integration.py— 9 passedpytest tests/core/test_erasure_service.py tests/core/test_erasure_saga.py tests/core/test_erasure_integration.py— 25 passedDesign review: three rubber-duck rounds with
opencode/space-bunny-free(hole → fix → re-verify); round-2 findings applied (kind-query logging, collect-then-mark in one txn, per-subject groups, manifest-on-close, negative tests).Manual:
slm restartclean (integrity ok);slm ops status→HEALTHYafter remediating pre-existing poisoned rows on the local 4.1.17 daemon (fix itself ships here, not yet installed locally).Reviewer Notes
git revertrestores old redrive; no schema migration.Fixes the
210a524c/ff8a4c44orphan-erase incident. See also #148 (adjacentops resolveCLI robustness, out of scope here).