EMR: labelled benchmark, lexical retrieval, and an age-free evidence gate - #8
Merged
Merged
Conversation
…gate A labelled benchmark (72 memories with realistic ages, 130 queries across keyword, natural-question, word-form, paraphrase, unrelated and near-miss categories), scored through emr_recall, the path the LLM route and MCP tools use. On it, current EMR recalled the right memory in the top 5 for 32% of answerable queries, 0% when the memory was over a month old, and answered 40% of unrelated questions. Causes and fixes: - Query alignment counted function words, so they diluted real questions and matched unrelated ones on "what"/"is". Terms are now stemmed content words, weighted by BM25 IDF over the ledger; an unseen query term weighs as much as the rarest real one, which reduces to plain overlap in a tiny ledger. - The abstention gate multiplied in decay and trajectory, so a draft fact could never be recalled after about three days. The gate is now relevance times provenance; decay still orders results, floored at 0.25 so it cannot erase a match. min_top_score moves to 0.08, above the 0.075 ceiling of a memory that shares no term with the query. - Memories sharing no term with the query entered STM on activation alone; they now enter only through graph expansion or as sticky prior STM. - emr_recall returned STM in fill order (activation per token), so short memories outranked better ones; the bundle is ordered by activation. - emr_eval's reinforcement probe needed an irrelevant memory already in the ranking; with unrelated memories kept out it found none and the positive-outcome gate failed on zero attempts. It now falls back to any irrelevant memory, which must stay out however often it is used. Result: 88% top-5 recall (85% on the half not used for tuning), 87% for memories over a month old, no unrelated question answered. tests/test_emr_bench holds those as a regression floor. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Live, the LLM route asked "Where does llm-gateway run, and on which port?". EMR recalled nothing, and the model invented a port. A labelled benchmark showed this is systematic: through
emr_recall, EMR found the right memory for 32% of answerable queries, 0% once the memory was over a month old, and returned memories for 40% of unrelated questions.The benchmark (
app/emr_bench.py,tests/fixtures/emr_bench/)Corpus: 72 memories, ages from 2 hours to 4 months, with a verified/draft mix, near-miss distractors and two conflict pairs.
Queries: 130 labelled queries:
kwnqmorphparanegnearFixed ages: ages are stamped relative to now, so decay sees the same ages on every run.
Real recall path: scored through
emr_recall, the path the LLM route and MCP tools use.emr_eval's contradiction, reinforcement and graph safety gates run alongside.Run:
python -m app.emr_bench --show-failures(about 3 s).What was wrong, and the fix
RECENCY_FLOOR = 0.25.min_top_scoremoves from 0.05 to 0.08, just above the 0.075 a memory sharing no query term can reach.emr_recallreturned STM in fill order, which is activation per token. The bundle is now ordered by activation.emr_eval's reinforcement probe needed an irrelevant memory already in the ranking. With (3) there's none, so the positive-outcome gate failed on zero attempts. It now falls back to any irrelevant memory, which must stay out however often it's used. The safety property itself is unchanged.Results
All 215 existing tests pass unchanged. New tests: 10 in
tests/test_emr_retrieval.py, plustests/test_emr_bench.py, which holds the scores above as a CI regression floor.For the reviewer
min_top_query_alignmentto 0.6 cuts near-miss answers to 20% but drops recall to 70%. That's left at 0.2 for now: a follow-up PR adds local embeddings, and this is its main test./exciteand the EMR→STM pipeline still usetheta_promote = 0.12, and activation includes the constant 0.2 trajectory term, so old memories still can't enter the working set through those paths. Only the recall path is fixed here.🤖 Generated with Claude Code