PARE combines structure-aware width pruning and timestep-conditioned adaptive depth routing to cut video diffusion compute while preserving generation quality.
NoRA evaluates visual first-person normative reasoning by requiring models to generate actions with fact-reason-action support graphs, revealing current VLMs struggle to bind correct justifications to actions.
MedVIGIL evaluates medical vision-language models under broken visual evidence via clinician-supervised probes, revealing a 14.1-point gap between top models and radiologist reliability.
ProCTI retrieves global dataset prototypes to augment local conditioning in diffusion-based time series imputation, improving accuracy under sparse or noisy missingness with theoretical guarantees.
TRL-Bench standardizes cross-paradigm evaluation of tabular encoders via shared representation-level probes, finding encoder quality is task-specific and best pipelines combine capability-matched specialists.
A diagnostic taxonomy maps AVLM failure signatures to targeted development interventions, enabling traceable industry-scale video moderation system improvements.
Contrastive identification and generation in the limit studies learning from unlabeled differing pairs, yielding geometric characterizations, a strict generation hierarchy, and robust corruption reversal via common crossing graphs.
EasyLens is a training-free plug-and-play module that amplifies subtle lesion representations in frozen medical vision-language models via prototype-based patch selection and morphology-guided residual enhancement, improving detection across datasets.
LOFT separates orthogonal PEFT subspaces from transformations to enable task-aware support selection, improving efficiency-performance trade-offs across language, vision, and reasoning tasks.
USAD improves adversarial detection via variance and perturbation covariance discrepancy statistics that capture global and local uncertainty patterns in adversarial examples.