Skip to content

EMR: labelled benchmark, lexical retrieval, and an age-free evidence gate - #8

Merged
warheart1984-ctrl merged 1 commit into
mainfrom
feat/emr-retrieval
Oct 2, 2026
Merged

warheart1984-ctrl merged 1 commit into
mainfrom
feat/emr-retrieval

Conversation

@warheart1984-ctrl

Copy link
Copy Markdown
Owner

Why

Live, the LLM route asked "Where does llm-gateway run, and on which port?". EMR recalled nothing, and the model invented a port. A labelled benchmark showed this is systematic: through emr_recall, EMR found the right memory for 32% of answerable queries, 0% once the memory was over a month old, and returned memories for 40% of unrelated questions.

The benchmark (app/emr_bench.py, tests/fixtures/emr_bench/)

  • Corpus: 72 memories, ages from 2 hours to 4 months, with a verified/draft mix, near-miss distractors and two conflict pairs.

  • Queries: 130 labelled queries:

    Prefix Category Cases
    kw Keyword 30
    nq Natural question 45
    morph Word forms 15
    para Paraphrase 20
    neg Unrelated, must abstain 10
    near Adjacent but unstored, must abstain 10
  • Fixed ages: ages are stamped relative to now, so decay sees the same ages on every run.

  • Real recall path: scored through emr_recall, the path the LLM route and MCP tools use. emr_eval's contradiction, reinforcement and graph safety gates run alongside.

  • Run: python -m app.emr_bench --show-failures (about 3 s).

What was wrong, and the fix

  1. Function words counted. "where / does / which" diluted real questions, and unrelated questions matched on "what / is / the". Query alignment now uses stemmed content words weighted by BM25 IDF over the ledger. An unseen query term weighs as much as the rarest real term, which reduces to plain overlap in a 1–2 memory ledger.
  2. Age erased memories. The abstention gate multiplied in decay and the constant trajectory term, so a draft fact became unrecallable after about 3 days. The gate is now relevance × provenance. Decay still orders results, but it's floored at RECENCY_FLOOR = 0.25. min_top_score moves from 0.05 to 0.08, just above the 0.075 a memory sharing no query term can reach.
  3. Unrelated memories rode along. Once one memory passed the gate, anything above activation 0.001 joined the bundle. Memories sharing no term with the query now enter only through graph expansion or as sticky prior STM.
  4. Shortness beat relevance. emr_recall returned STM in fill order, which is activation per token. The bundle is now ordered by activation.
  5. The eval harness's distractor probe. emr_eval's reinforcement probe needed an irrelevant memory already in the ranking. With (3) there's none, so the positive-outcome gate failed on zero attempts. It now falls back to any irrelevant memory, which must stay out however often it's used. The safety property itself is unchanged.

Results

Before After
Top-5 recall, answerable 0.32 0.88 (0.85 on the half not used for tuning)
Top-1 0.25 0.67
Memory 1–4 weeks old 0.16 0.91
Memory over 1 month old 0.00 0.87
Answerable but abstained 0.36 0.05
Unrelated question answered 0.40 0.00
Paraphrase 0.05 0.45
Near-miss answered (should abstain) 0.60 1.00 ⚠️

All 215 existing tests pass unchanged. New tests: 10 in tests/test_emr_retrieval.py, plus tests/test_emr_bench.py, which holds the scores above as a CI regression floor.

For the reviewer

  • Near-misses regressed. Before, they mostly abstained by accident, because decay suppressed the memory. Now "sister's birthday" returns the brother's birthday and the sister's city. Lexical scoring can't separate these without losing real recall. Raising min_top_query_alignment to 0.6 cuts near-miss answers to 20% but drops recall to 70%. That's left at 0.2 for now: a follow-up PR adds local embeddings, and this is its main test.
  • Gate semantics change. Age no longer counts as evidence against a memory. This is deliberate: a year-old verified fact is no less true. Tasks may deserve an exception, since they go stale; open to that.
  • Not changed: /excite and the EMR→STM pipeline still use theta_promote = 0.12, and activation includes the constant 0.2 trajectory term, so old memories still can't enter the working set through those paths. Only the recall path is fixed here.
  • Synthetic benchmark. The benchmark is hand-written, not a real ledger. A second check on a real ledger, with labels, should come before trusting the numbers fully.

🤖 Generated with Claude Code

…gate

A labelled benchmark (72 memories with realistic ages, 130 queries across
keyword, natural-question, word-form, paraphrase, unrelated and near-miss
categories), scored through emr_recall, the path the LLM route and MCP tools
use. On it, current EMR recalled the right memory in the top 5 for 32% of
answerable queries, 0% when the memory was over a month old, and answered
40% of unrelated questions.

Causes and fixes:
- Query alignment counted function words, so they diluted real questions and
  matched unrelated ones on "what"/"is". Terms are now stemmed content words,
  weighted by BM25 IDF over the ledger; an unseen query term weighs as much
  as the rarest real one, which reduces to plain overlap in a tiny ledger.
- The abstention gate multiplied in decay and trajectory, so a draft fact
  could never be recalled after about three days. The gate is now relevance
  times provenance; decay still orders results, floored at 0.25 so it cannot
  erase a match. min_top_score moves to 0.08, above the 0.075 ceiling of a
  memory that shares no term with the query.
- Memories sharing no term with the query entered STM on activation alone;
  they now enter only through graph expansion or as sticky prior STM.
- emr_recall returned STM in fill order (activation per token), so short
  memories outranked better ones; the bundle is ordered by activation.
- emr_eval's reinforcement probe needed an irrelevant memory already in the
  ranking; with unrelated memories kept out it found none and the
  positive-outcome gate failed on zero attempts. It now falls back to any
  irrelevant memory, which must stay out however often it is used.

Result: 88% top-5 recall (85% on the half not used for tuning), 87% for
memories over a month old, no unrelated question answered. tests/test_emr_bench
holds those as a regression floor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@warheart1984-ctrl
warheart1984-ctrl merged commit c207247 into main Oct 2, 2026
2 checks passed
@warheart1984-ctrl
warheart1984-ctrl deleted the feat/emr-retrieval branch October 2, 2026 02:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant