VIGIL decouples world-state completion from terminal commitment in embodied agents, revealing that comparable execution yields up to 19.7 pp differences in correct episode termination.
PhysVista benchmarks physical intelligence in vision-language models via a perception-reasoning-assessment loop, exposing major gaps in physical reasoning and plausibility assessment.
BitDance is an autoregressive image generator that predicts binary visual tokens via a diffusion head and next-patch decoding, achieving state-of-the-art FID with far fewer parameters and much faster inference.
SMTL replaces sequential reasoning with parallel evidence acquisition for efficient long-horizon agentic search, achieving state-of-the-art results on multiple benchmarks with far fewer reasoning steps.
IPIBench evaluates interactive proactive intelligence of MLLMs on continuous video streams, revealing unstable proactive triggering and weak reactive-proactive coordination, while IPI-Agent improves both via temporal gating.
xMemory decouples agent memories into reusable components before aggregating them hierarchically, improving retrieval quality and token efficiency over flat RAG.
SplitMoE replaces uniform token-wise routing with split semantic and generic experts, improving video diffusion convergence, routing coherence, and generation quality over load-balanced MoEs.
VisInteract introduces interactive text-to-visualization with imperfect queries via VisInteract-Bench and Vis-MCTS, boosting success by over 13% versus interactive baselines.
Edit-R2 uses reinforcement learning to reconstruct session intent and jointly optimize reasoning and generation for multi-turn image editing. It improves instruction following and consistency over accumulated constraints on the MICE-Bench benchmark.
HOMIE unifies inter- and intra-subject video personalization via multimodal guidance and reference embeddings, achieving state-of-the-art human-object interaction fidelity.
Generation Navigator is a state-aware multi-turn text-to-image agent that learns to steer generation via trajectory-level reinforcement learning, achieving a 0.90 WISE score and 79.06% reasoning accuracy.