EMR: optional local embeddings for semantic recall - #10
Merged
Merged
Conversation
Lexical alignment cannot see that "what is the pet called" asks about "Jon has a dog named Bruno". app/emr_embed.py adds an optional semantic signal: a small local model (fastembed, BAAI/bge-small-en-v1.5, ONNX on CPU, 65 MB, ~13 ms per sentence) embeds the query and each memory, and Q becomes the stronger of lexical and semantic alignment. Cosine maps onto Q's scale between SIM_FLOOR 0.65 and SIM_FULL 0.85, calibrated on half the benchmark. Off unless JARVIS_EMR_EMBEDDINGS=1 and the `embed` extra is installed; if the model cannot load, EMR stays lexical. Memory vectors are cached by content hash and model in JARVIS_EMR_EMBED_CACHE. ActivationBreakdown gains Q_lex and Q_sem so a recall can be audited. Benchmark with embeddings on, against the lexical stage: top-5 0.88 -> 0.95, top-1 0.67 -> 0.74, paraphrases 0.45 -> 0.70, memories over a month old 0.87 -> 0.97, unrelated questions answered stays 0.00. `--embeddings` runs it; the embedding floors in tests/test_emr_bench skip without the extra. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The previous commit rewrote app/emr.py with LF endings, so its diff showed every line as changed. The file is CRLF on main; content is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
app/emr_embed.py, an optional semantic signal for EMR query alignment. A small local model embeds the query and each memory, andQ = max(Q_lex, Q_sem).BAAI/bge-small-en-v1.5. It runs as ONNX on the CPU with no torch, takes 65 MB on disk, and embeds about 13 ms per sentence on a 4-core box.SIM_FLOOR = 0.65(no signal) andSIM_FULL = 0.85(full alignment).JARVIS_EMR_EMBED_CACHE, so each memory is embedded once.ActivationBreakdowngainsQ_lexandQ_sem, so any recall can be audited.pip install ".[embed]"andJARVIS_EMR_EMBEDDINGS=1. If the model can't load (missing extra, offline first run, bad name), EMR stays lexical and recall doesn't fail.python -m app.emr_bench --embeddingsscores the benchmark with embeddings on.Results (same 130-case benchmark as #8)
Calibration. I picked the floor and ceiling on the odd-numbered half of the cases and checked them on the even half. The dev-best setting (0.60–0.85) dropped on the held-out half: paraphrases fell from 1.00 to 0.60, with only 10 paraphrases per half. 0.65–0.85 was nearly tied on dev and held up on held-out (0.94 top-5, 0.76 top-1), so that's what ships. Disclosure: I looked at the held-out half for four configurations, so it's not perfectly clean.
Near-misses, measured end to end. Embeddings don't separate near-misses: "sister's birthday" scores 0.79 against "brother's birthday", higher than a true paraphrase. So I measured what matters, the answer. All 10 near-miss questions went through recall plus
groq/gpt-oss-20bvia llm-gateway, with the recalled memories as context. 9 of 10 answered correctly ("not on record"). The one miss divided the Jarvis tenant's daily budget by 24 for "spend per hour on OpenAI", ignoring "OpenAI". The paraphrases tried (truck, coffee, pet) and the original failing question ("Where does llm-gateway run, and on which port?" now returns127.0.0.1:8080with a citation) were all correct. The retrieval-level near-miss number overstates the harm: returning "Jon has a dog named Bruno" for "what's the cat's name?" is the context that lets the model say there's no cat on record.Tests
tests/test_emr_embed.pyuses a fake model, so there's no download. It covers: off by default, an unloadable model falls back to lexical, the cosine mapping, a paraphrase with no shared words being recalled, Q as the max of the two signals, and the vector cache (per content hash, and invalidated by a model change).tests/test_emr_bench.pyholds the numbers above as a regression floor, and skips without the extra.main(AMUL LLM: llm-gateway backend, API keys, and HTTP routes #7 + EMR: labelled benchmark, lexical retrieval, and an age-free evidence gate #8): 254 passed with the extra installed (3.12). On Python 3.11 with.[dev]only (as CI runs it), 252 passed and 2 skipped.For the reviewer
JARVIS_EMR_EMBED_MODEL_DIR(defaultdata/emr-embed-model/, now git-ignored). After that it runs offline. The Docker image doesn't install the extra.🤖 Generated with Claude Code