Conversation
The control arm measures what retrieval actually contributes by passing the
entire untruncated history of each instance directly to the reader model.
The benchmark harness was extended with the `full_history` arm to bypass the
retrieval step and pass all sessions in date order.
Results (n=500, LongMemEval-S):
- `full_history` accuracy: 94.60%
- `rerank12` accuracy (shipped retrieval): 94.20%
The 0.4% difference is statistical noise (2 questions out of 500).
Retrieval reduces token usage massively (from ~125,000 down to under
5,000 per request) while preserving the accuracy of the full context.
docs/benchmarks/E16-full-history-control.md details the setup, smoke
run, and final results. docs/benchmarks/E9-peer-protocols.md adds the
control arm to Table 3. docs/RESULTS.md closes the "Net saving versus
no memory — UNMEASURED" gap with the empirical savings.
Signed-off-by: doronp <3586743+doronp@users.noreply.github.com>
…by warm-cache rerun)
…table 94.60 warm-cache)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds a
full_historycontrol arm tobench/qa.py(QA-pipeline-level, not retrieval-metric-level): bypasses retrieval entirely and passes every session id, date-sorted, into the same reader and judge as the published arms. Prompts untouched (sha256 pins green),bench/peer.pyunaffected (QA_ARMSis qa-local).Result (n=500, LongMemEval-S, reader=judge=gemini-3.1-pro-preview, temp 0)
rerank12(published)full_history(control)Identical within sampling error (n=500). Pre-registered read: retrieval saves ~95% of reader tokens (~125k → ~5k per question) with no measurable accuracy penalty. Closes the "Net saving versus no memory — UNMEASURED" dashboard gap in
docs/RESULTS.md.Verification
bench/fetch_longmemeval.sh.bench/test_qa.py: 10 passed, incl. new tests that the arm passes all session ids in date order and never calls_retrieve.docs/benchmarks/E16-full-history-control.md.Cost: ~62.5M input tokens, ~12 min at 64 workers on Vertex (credits project).