ACT*ONOMY introduces a three-level taxonomy of 10 actions and 46 subactions for describing autonomous agent behavior at runtime, plus an open repository and automated analysis pipeline that compares behavioral profiles and surfaces failure patterns.
CausalSpatial benchmarks object-centric causal spatial reasoning, revealing MLLMs score 54% versus human 84% due to ungrounded textual reasoning, fixed by video-simulation framework COW.
AgentOdyssey generates open-ended long-horizon text games to evaluate test-time continual learning, finding top agents far below human performance despite scaling with model strength.
DARLING uses a learned partition function to jointly optimize language model response quality and semantic diversity via reinforcement learning, improving both quality and novelty across creative and math benchmarks.
RATS decomposes vision transformers' classification token into learnable register tokens that spontaneously specialize into object parts, improving segmentation by up to 12 mIoU.
Proposed E-P-R framework diagnoses AI agents consuming conflicting memory via entry-propagation-recovery, finding a compliance trap where early adoption collapses success.