MemForest partitions agent memory into event trees and progressively merges redundant nodes to cut storage and retrieval costs while preserving nearly all performance.
MM-IssueLoc benchmarks multimodal repository-level issue localization using visual evidence across 652 instances, showing current systems achieve under 39% file accuracy and text-only scores do not transfer.
L2P transfers pre-trained latent diffusion models to pixel space via frozen intermediate layers and synthetic data, enabling efficient 4K generation with near-source performance.
ViDiC introduces a video difference captioning task and ViDiC-1K benchmark that reveals large multimodal models struggle with fine-grained comparative video perception.
ExpLang improves LLM reasoning via on-policy multilingual thinking language selection during RL, outperforming English-only training and extending exploration with diverse language preferences.
ABPR couples LLMs with Prolog meta-interpreters to debug programs via proof-tree analysis, achieving up to 98.33% Pass@2 on ARC-AGI-2 and extending to relational reasoning benchmarks.
Linguistic Trajectory Encoding compresses long-horizon object motion into hybrid language-spatial-visual timelines, outperforming baselines on multi-day spatial memory benchmarks with high compression and sub-second queries.
EcoGym benchmarks long-horizon LLM economic planning across open-source environments, revealing no single model dominates and exposing strategic and execution suboptimalities.
Compressed distributed online convex optimization achieves optimal regret via error feedback and online compression, with applications to distributed non-smooth optimization.
TerminalWorld automatically builds terminal benchmarks from wild recordings, yielding 1,530 tasks where top agents achieve only 62.5% success with weak correlation to expert benchmarks.
WebNavigator overcomes topological blindness via interaction graphs to turn web navigation into deterministic retrieval and pathfinding, doubling multi-site success on WebArena.