ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge
ARK introduces a dual-axis multimodal retrieval benchmark spanning knowledge domains and reasoning skills, revealing persistent bottlenecks in fine-grained visual and spatial reasoning.
Published 2026Sydney Poster Session 3 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

Only vote on papers you've read. Sign in with GitHub to vote.
Abstract
Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To address this gap, we introduce ARK, a benchmark designed to analyze multimodal retrieval from two complementary perspectives: (i) knowledge domains (five domains with 17 subtypes), which characterize the content and expertise retrieval relies on, and (ii) reasoning skills (six categories), which characterize the type of inference over multimodal evidence required to identify the correct candidate. Specifically, ARK evaluates retrieval with both unimodal and multimodal queries and candidates, covering 16 heterogeneous visual data types. To avoid shortcut matching during evaluation, most queries are paired with targeted hard negatives that require multi-step reasoning. We evaluate 25 representative text-based and multimodal retrievers and observe a pronounced gap between knowledge- and reasoning-intensive retrieval, with fine-grained visual and spatial reasoning as persistent bottlenecks. We further show that enhancements such as re-ranking, rewriting, and agentic retrieval yield consistent gains, but substantial headroom remains.