Memory-R2 proposes LoGo-GRPO to enable fair credit assignment for memory-augmented LLM agents across long multi-session horizons via local rerollouts and shared-parameter co-learning.
Symb-xMIL quantifies alignment between MIL predictions and human-readable logical rules to expose decision patterns, recover ground-truth rules, and refine survival stratification beyond HPV status.