PARE combines structure-aware width pruning and timestep-conditioned adaptive depth routing to cut video diffusion compute while preserving generation quality.
NoRA evaluates visual first-person normative reasoning by requiring models to generate actions with fact-reason-action support graphs, revealing current VLMs struggle to bind correct justifications to actions.
MedVIGIL evaluates medical vision-language models under broken visual evidence via clinician-supervised probes, revealing a 14.1-point gap between top models and radiologist reliability.