Good Papers

RelayLLM: Efficient Reasoning via Collaborative Decoding

RelayLLM enables small language models to dynamically invoke large models for critical reasoning tokens via collaborative decoding, reducing costs by 98.2% while achieving 49.52% accuracy.

Chengsong Huang, Tong Zheng, Langlin Huang, Jinyuan Li, Haolin Liu, Jiaxin Huang

Published Jan 8, 2026▲ 30 on Hugging FaceCode ★ 41arXiv ↗

80%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel12/20reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
RelayLLM delivers a genuinely striking token-level relay that slashes LLM usage to near 1% and narrows the reasoning gap, yet its benchmark fantasy collapses without open code, wall-clock latency proofs, and clear failure modes for bad…

Abstract

Large Language Models (LLMs) for complex reasoning is often hindered by high computational costs and latency, while resource-efficient Small Language Models (SLMs) typically lack the necessary reasoning capacity. Existing collaborative approaches, such as cascading or routing, operate at a coarse granularity by offloading entire queries to LLMs, resulting in significant computational waste when the SLM is capable of handling the majority of reasoning steps. To address this, we propose RelayLLM, a novel framework for efficient reasoning via token-level collaborative decoding. Unlike routers, RelayLLM empowers the SLM to act as an active controller that dynamically invokes the LLM only for critical tokens via a special command, effectively "relaying" the generation process. We introduce a two-stage training framework, including warm-up and Group Relative Policy Optimization (GRPO) to teach the model to balance independence with strategic help-seeking. Empirical results across six benchmarks demonstrate that RelayLLM achieves an average accuracy of 49.52%, effectively bridging the performance gap between the two models. Notably, this is achieved by invoking the LLM for only 1.07% of the total generated tokens, offering a 98.2% cost reduction compared to performance-matched random routers.