Good Papers

When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model

Raw-turn selection via a typed decision model matches LLM extraction at tight budgets but falls behind at generous budgets, explaining conflicting memory results.

Rishabh Sharma, Rishika Lall

Published Sep 28, 2026▲ 11 on Hugging FaceCode ★ 1arXiv ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Pre-registered results show raw-turn selection is non-inferior to extraction at tight budgets with a 3,061x write cost advantage, though extraction wins generously and Jev's typed model is appendix-buried.

Abstract

Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.