SIEVES improves selective prediction for visual question answering by scoring visual evidence quality, boosting out-of-distribution coverage up to three times across open and closed models without requiring internal weights.
ZendoWorld evaluates AI agents on active visual rule induction and finds high prediction accuracy does not imply rule recovery, with VLM agents proposing near-uninformative experiments.
GeRo enables vision-language-action models to generate language-grounded future traffic scenes via autoregressive rollouts, improving Bench2Drive driving scores by 15.7 and success rates by 26.2.
NAUTILUS converts a single prompt into robot learning workflows via plug-and-play agent skills, typed contracts, and automated validation, reducing cross-family engineering overhead.