Good Papers

Capability-Driven Self-Evolution of Agent Memory

PrisMem drives agent memory self-evolution via capability-specific guidance, dependency-aware selection, and trace-guided integration, outperforming baselines by up to 10.54 points on million-token benchmarks.

Yaoqi Chen, Yuru Feng, Qianxi Zhang, Baotong Lu, Jianan Lu, Zhirui Wang, Shusen Xu, Zewen Jin, Zengzhong Li, Cheng Li, Qi Chen

Published Oct 5, 2026▲ 6 on Hugging FacearXiv ↗

80%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel12/20reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Capability-driven self-evolution offers a sharp diagnosis of mixed-feedback stagnation and delivers striking benchmark gains, yet its ten-point lift rests on synthetic long-context tests with unclear baseline rigor, unverified scalability, and hidden trace-curation overhead.

Abstract

Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and judging progress by overall performance. This can obscure optimization directions and hide capability-specific gains offset by regressions elsewhere, leaving promising directions underexplored. We introduce capability-driven evolution, which extends search guidance from overall performance to individual capability dimensions, preserving promising revisions and expanding exploration beyond the boundaries of holistic evolution. We propose PrisMem, which uses dependency-aware capability selection to prioritize targets with potential cross-capability benefits and history-guided diagnosis to refine capability specialists. Trace-guided integration compares evaluated programs on paired differential cases, using their behavioral differences to consolidate complementary gains into a unified memory program. Experiments show that PrisMem outperforms the strongest baselines by 10.54 and 7.83 percentage points on BEAM-1M and LongMemEval-M, respectively, demonstrating its effectiveness on million-token histories.