A reviewer noted that Honcho's benchmark post (linked from our comparison table) reports Gemini 3 Pro alone — full history, no memory — at 92.0 on LongMemEval-S, i.e. 2.2 points below our 94.2 and plausibly within sampling error on 500 questions. Their reader is a version older and their judge was GPT-4o, so the numbers are not directly comparable.
Proposed control, credit to the reviewer: run a no-retrieval arm in our harness — same reader, same judge, full untruncated history instead of retrieval.
Interpretation:
- If it lands within noise of 94.2: on this benchmark retrieval saves tokens but adds no measurable accuracy; the 94.2 reflects the reader more than the memory.
- If it lands clearly below: we have a clean measurement of what retrieval adds, isolated from the reader.
Either way the result gets published on the dashboard alongside the existing arms, and the compaction result is unaffected (there the evidence has left the window; live window scores 0.0).
Related existing gap already on the dashboard: 'Net saving versus no memory — UNMEASURED.'
A reviewer noted that Honcho's benchmark post (linked from our comparison table) reports Gemini 3 Pro alone — full history, no memory — at 92.0 on LongMemEval-S, i.e. 2.2 points below our 94.2 and plausibly within sampling error on 500 questions. Their reader is a version older and their judge was GPT-4o, so the numbers are not directly comparable.
Proposed control, credit to the reviewer: run a no-retrieval arm in our harness — same reader, same judge, full untruncated history instead of retrieval.
Interpretation:
Either way the result gets published on the dashboard alongside the existing arms, and the compaction result is unaffected (there the evidence has left the window; live window scores 0.0).
Related existing gap already on the dashboard: 'Net saving versus no memory — UNMEASURED.'