76%Highly rated
?Highly ratedVote to see the score

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents
Metacognitive Memory Policy Optimization uses belief entropy to penalize uncertain intermediate summaries, improving long-horizon LLM agent reasoning and retaining 97.1% performance at 1.75M-token contexts.
Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026 · ▲ 7 on Hugging Face
– ReadersNo votes yet
10/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5