Skip to content

bench: full-history no-retrieval control arm on LongMemEval-S (closes #16) - #17

Open
doronp wants to merge 4 commits into
mainfrom
bench/full-history-control
Open

doronp wants to merge 4 commits into
mainfrom
bench/full-history-control

Conversation

@doronp

@doronp doronp commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

What

Adds a full_history control arm to bench/qa.py (QA-pipeline-level, not retrieval-metric-level): bypasses retrieval entirely and passes every session id, date-sorted, into the same reader and judge as the published arms. Prompts untouched (sha256 pins green), bench/peer.py unaffected (QA_ARMS is qa-local).

Result (n=500, LongMemEval-S, reader=judge=gemini-3.1-pro-preview, temp 0)

arm accuracy
rerank12 (published) 94.20
full_history (control) 94.60

Identical within sampling error (n=500). Pre-registered read: retrieval saves ~95% of reader tokens (~125k → ~5k per question) with no measurable accuracy penalty. Closes the "Net saving versus no memory — UNMEASURED" dashboard gap in docs/RESULTS.md.

Verification

  • Orchestrator independently reproduced the published table with the final code (0.9460 overall, every per-type row identical; an earlier 64-worker run scored 93.80 with 6/500 flips — see E16 doc rerun-variance note). A second fully cache-warm rerun reproduced 0.9460 with zero new API calls.
  • Dataset sha256 matches the pinned value in bench/fetch_longmemeval.sh.
  • bench/test_qa.py: 10 passed, incl. new tests that the arm passes all session ids in date order and never calls _retrieve.
  • Full details: docs/benchmarks/E16-full-history-control.md.

Cost: ~62.5M input tokens, ~12 min at 64 workers on Vertex (credits project).

doronp added 3 commits October 3, 2026 15:03
    The control arm measures what retrieval actually contributes by passing the
    entire untruncated history of each instance directly to the reader model.
    The benchmark harness was extended with the `full_history` arm to bypass the
    retrieval step and pass all sessions in date order.

    Results (n=500, LongMemEval-S):
    - `full_history` accuracy: 94.60%
    - `rerank12` accuracy (shipped retrieval): 94.20%

    The 0.4% difference is statistical noise (2 questions out of 500).
    Retrieval reduces token usage massively (from ~125,000 down to under
    5,000 per request) while preserving the accuracy of the full context.

    docs/benchmarks/E16-full-history-control.md details the setup, smoke
    run, and final results. docs/benchmarks/E9-peer-protocols.md adds the
    control arm to Table 3. docs/RESULTS.md closes the "Net saving versus
    no memory — UNMEASURED" gap with the empirical savings.

    Signed-off-by: doronp <3586743+doronp@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant