Good Papers

M$^\star$: Every Task Deserves Its Own Memory Harness

M* evolves task-specific memory programs via reflective code search to outperform fixed-memory agents across diverse benchmarks. Evolved harnesses develop structurally distinct mechanisms per domain, showing specialization beats general-purpose memory.

Wenbo Pan, Shujie LIU, Xiangyang Zhou, Xianlong Wang, Shiwei Zhang, Bingjing Xu, Wanlu Shi, Xiaohua Jia

Published 2026Sydney Poster Session 2 · Tue, Dec 8, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

80%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel12/20reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Large language model agents rely on specialized memory systems to accumulate and reuse knowledge during extended interactions. Recent architectures typically adopt a fixed memory design tailored to specific domains, such as semantic retrieval for conversations or skills reused for coding. However, a memory system optimized for one purpose frequently fails to transfer to others. To address this limitation, we introduce M$^\star$, a method that automatically discovers task-optimized memory harnesses through executable program evolution. Specifically, M$^\star$ models an agent memory system as a memory program written in Python. This program encapsulates the data Schema, the storage Logic, and the agent workflow Instructions. We optimize these components jointly using a reflective code evolution method; this approach employs a population-based search strategy and analyzes evaluation failures to iteratively refine the candidate programs. We evaluate M$^\star$ on four distinct benchmarks spanning conversation, embodied planning, and expert reasoning. Our results demonstrate that M$^\star$ improves performance over existing fixed-memory baselines robustly across all evaluated tasks. Furthermore, the evolved memory programs exhibit structurally distinct processing mechanisms for each domain. This finding indicates that specializing the memory mechanism for a given task explores a broad design space and provides a superior solution compared to general-purpose memory paradigms.