Good Papers

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents

Metacognitive Memory Policy Optimization uses belief entropy to penalize uncertain intermediate summaries, improving long-horizon LLM agent reasoning and retaining 97.1% performance at 1.75M-token contexts.

Ziyan Liu, Zhezheng Hao, Yeqiu Chen, Hong Wang, JingRen Hou, Ruiyi Ding, Yongkang Yang, Wence Ji, Wei Xia, Feng Liu

Published 2026Paris Poster Session 2 · Wed, Dec 9, 5:00 PM–7:00 PM local time · Paris Poster Hall▲ 7 on Hugging FacearXiv ↗OpenReview ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Memory-augmented LLM agents tackle complex long-horizon tasks by recursively summarizing interaction trajectories into compact memory. However, existing approaches typically train these memory policies using outcome-based reinforcement learning, failing to localize where intermediate memory quality degrades. As interactions unfold, ambiguous recursive summaries progressively discard task-relevant information and introduce semantic noise. This exacerbates belief deviation, obscuring the agent's estimate of the latent task state and ultimately derailing long-horizon reasoning. We therefore argue that memory optimization should focus not merely on trajectory-level success, but on the clarity of the belief induced by intermediate summaries. To this end, we introduce Belief Entropy, a self-supervised proxy that probes how uncertain the model remains about the latent task state given its current memory. Based on this proxy, we propose Metacognitive Memory Policy Optimization (MMPO). Instead of relying only on sparse outcome-based signals, MMPO provides fine-grained, memory-specific supervision via explicitly penalizing summaries that induce high epistemic uncertainty. Experiments show that MMPO consistently outperforms existing methods on diverse long-horizon tasks, maintaining 97.1% performance even when scaled to 1.75M-token contexts.