Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents
Selection-based Structured Reasoning replaces open-ended reasoning with selection among reusable candidates, cutting per-turn latency over 90% while matching leading small-model search agents' success rates.
Published Oct 1, 2026▲ 3 on Hugging FaceCodearXiv ↗
Only vote on papers you've read. Sign in with GitHub to vote.
SSR delivers striking latency reductions and competitive search success on small models by replacing open-ended reasoning with selection, though it lacks ablations on candidate design and cross-domain transfer.
Abstract
Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.