WorldMemArena evaluates multimodal agent memory through an action-world loop, showing writing and storage improvements do not guarantee performance and harness-based memory remains costly and unreliable.
OmniSpace improves autonomous vehicle MLLM spatial reasoning via camera pose injection, multi-view epipolar attention, and 3D geometric distillation without auxiliary 3D models, surpassing existing methods across planning, risk detection, and language benchmarks.
Analysis of 20 LLMs finds confidence reflects shared item difficulty rather than individuated self-assessment, showing no evidence of functional metacognition.
For any k>2, (k+1)/(k+2)-EFkX allocations always exist and are computable in polynomial time, yielding 3/4-EF2X for any number of agents and 2/3-EF X for eight agents.
DoAtlas-1 introduces causal compilation to convert medical evidence into executable causal estimands, achieving 98.5% canonicalization accuracy and 80.5% query executability across 1,445 effect kernels.
NeuroDoc introduces a rulebook-guided task specification layer that standardizes EEG benchmarks into 53 reviewed entries with 245 executable task definitions across four model backbones.
BusterX introduces GenBuster-200K, GenBuster-Bench, and an MLLM baseline that detects AI-generated video via reasoning chains, outperforming leading models in accuracy and explanation quality.
DriveSpatial benchmarks vision-language models' spatiotemporal autonomous driving intelligence, finding a 28.4-point human gap with cognitive scene construction as the key bottleneck.