VIGIL decouples world-state completion from terminal commitment in embodied agents, revealing that comparable execution yields up to 19.7 pp differences in correct episode termination.
PhysVista benchmarks physical intelligence in vision-language models via a perception-reasoning-assessment loop, exposing major gaps in physical reasoning and plausibility assessment.
BitDance is an autoregressive image generator that predicts binary visual tokens via a diffusion head and next-patch decoding, achieving state-of-the-art FID with far fewer parameters and much faster inference.
SMTL replaces sequential reasoning with parallel evidence acquisition for efficient long-horizon agentic search, achieving state-of-the-art results on multiple benchmarks with far fewer reasoning steps.
IPIBench evaluates interactive proactive intelligence of MLLMs on continuous video streams, revealing unstable proactive triggering and weak reactive-proactive coordination, while IPI-Agent improves both via temporal gating.
xMemory decouples agent memories into reusable components before aggregating them hierarchically, improving retrieval quality and token efficiency over flat RAG.
SplitMoE replaces uniform token-wise routing with split semantic and generic experts, improving video diffusion convergence, routing coherence, and generation quality over load-balanced MoEs.
VisInteract introduces interactive text-to-visualization with imperfect queries via VisInteract-Bench and Vis-MCTS, boosting success by over 13% versus interactive baselines.
Edit-R2 uses reinforcement learning to reconstruct session intent and jointly optimize reasoning and generation for multi-turn image editing. It improves instruction following and consistency over accumulated constraints on the MICE-Bench benchmark.
HOMIE unifies inter- and intra-subject video personalization via multimodal guidance and reference embeddings, achieving state-of-the-art human-object interaction fidelity.
Generation Navigator is a state-aware multi-turn text-to-image agent that learns to steer generation via trajectory-level reinforcement learning, achieving a 0.90 WISE score and 79.06% reasoning accuracy.
A diagnostic taxonomy maps AVLM failure signatures to targeted development interventions, enabling traceable industry-scale video moderation system improvements.
BusterX introduces GenBuster-200K, GenBuster-Bench, and an MLLM baseline that detects AI-generated video via reasoning chains, outperforming leading models in accuracy and explanation quality.
cIPO aligns text-to-video diffusion by deriving implicit preferences from reconstruction errors and concentrating optimization on high-error temporal segments to fix sparse artifacts.
SkipSR accelerates diffusion video super-resolution by skipping low-detail regions identified in low-resolution inputs, achieving up to 60% faster latency without quality loss.
ActWorld extends interactive world models to object interaction via a 100K dataset and hierarchical action-aware memory, improving fidelity over navigation-only baselines.
IRPO applies GRPO post-training to image restoration via selective hard-sample data and multi-component rewards, improving in-domain accuracy by 0.93 dB and OOD generalization by 3.43 dB.
EasyLens is a training-free plug-and-play module that amplifies subtle lesion representations in frozen medical vision-language models via prototype-based patch selection and morphology-guided residual enhancement, improving detection across datasets.
NLAC trains LLM agents with a natural-language generative critic for off-policy learning, yielding richer feedback and more stable, data-efficient training than policy gradients in long-horizon tasks.
SPIRAL uses sequential planning and reflective agents in a closed loop to generate long-horizon action-conditioned videos with iterative refinement and self-evolving post-training.
A generic patch Transformer achieves state-of-the-art zero-shot time series forecasting via simple training, with scaling and data ablations isolating key performance drivers.