Skip to content

Bench: full-history no-retrieval control arm on LongMemEval-S #16

Description

@doronp

A reviewer noted that Honcho's benchmark post (linked from our comparison table) reports Gemini 3 Pro alone — full history, no memory — at 92.0 on LongMemEval-S, i.e. 2.2 points below our 94.2 and plausibly within sampling error on 500 questions. Their reader is a version older and their judge was GPT-4o, so the numbers are not directly comparable.

Proposed control, credit to the reviewer: run a no-retrieval arm in our harness — same reader, same judge, full untruncated history instead of retrieval.

Interpretation:

  • If it lands within noise of 94.2: on this benchmark retrieval saves tokens but adds no measurable accuracy; the 94.2 reflects the reader more than the memory.
  • If it lands clearly below: we have a clean measurement of what retrieval adds, isolated from the reader.

Either way the result gets published on the dashboard alongside the existing arms, and the compaction result is unaffected (there the evidence has left the window; live window scores 0.0).

Related existing gap already on the dashboard: 'Net saving versus no memory — UNMEASURED.'

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions