Good Papers

REMORY: Learning Residual Memory for Context Compaction

REMORY adds learnable soft memory tokens to summaries to approximate full-context outputs with minimal positions, improving long-horizon agent attribution, coverage, and tool-use accuracy.

Hanchen Xia, Baoyou Chen, Yutang Ge, Naihao Deng, Senqiao Yang, Zilong Dong, Weihao Yuan, Siyu Zhu

Published Oct 8, 2026▲ 1 on Hugging FaceCodearXiv ↗

78%
OverallHighly rated
?
OverallHighly ratedVote to see the score
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel11/20reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce with the full history. The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension. On SummHay, REMORY improves source attribution at nearly unchanged insight coverage and approaches the full-context joint score using only 5.2% of the input positions. Across long-horizon agent benchmarks, Qwen3.8-27B and GLM-5.3-Flash show consistent gains with residual memory. Both models also exhibit substantially fewer repeated tool outputs and tool errors on BrowseComp and Terminal-Bench 2.1.