CausalSpatial benchmarks object-centric causal spatial reasoning, revealing MLLMs score 54% versus human 84% due to ungrounded textual reasoning, fixed by video-simulation framework COW.
AgentOdyssey generates open-ended long-horizon text games to evaluate test-time continual learning, finding top agents far below human performance despite scaling with model strength.
RISE sketches LLM output-layer influence hotspots into compressed dual-channel sketches, reducing storage up to 112x versus gradient methods while scaling to 32B parameters for attribution and data valuation.
AnyHand provides 6.6M synthetic RGB-D hand images with occlusions and aligned depth, significantly improving 3D hand pose estimation benchmarks and showing data diversity rivals scale.