When One LLM Drools, Multi-LLM Collaboration Rules
Multi-LLM collaboration outperforms single LLM reasoning on tasks where individual models fail, demonstrating collective rule over solo drooling.
Published 20261 citationPaper ↗

67%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel2/20reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
The paper makes multi-LLM collaboration a compelling framework for analyzing error modes, but without code, latency benchmarks, reproducible baselines, or clarity on whether gains come from disagreement rather than compute, its practical impact remains unproven.
Abstract
Shangbin Feng, Wenxuan Ding, Alisa Liu, Zifeng Wang, Weijia Shi, Yike Wang, Shannon Zejiang Shen, Xiaochuang Han, Hunter Lang, Chen-Yu Lee, Tomas Pfister, Yejin Choi, Yulia Tsvetkov. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.