Good Papers

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

MiroThinker introduces interaction scaling to train open-source research agents for deeper tool use, achieving up to 81.9% on GAIA and rivaling commercial models.

MiroMind Team, Bai, Song, Lidong Bing, Chen, Carson, Guanzheng Chen, Yuntao Chen, Z. Chen, Ziyi Chen, Yifeng Dai, Dong, Xuan, Dou, Wenhan, Yue Deng, Fu, Yunjie, Ge, Junqi, Chenxia Han, Tammy T. Huang, Zhenhang Huang, Jiao, Jerry, Jiang, Shilei, Jiao, Tianyu, Jian, Xiaoqi, Lei Lei, Ruilin Li, Luo, Gen, Tingan Li, Xiang Lin, Ziyuan Liu, Zhiqi Li, Jie Ni, Qiang Ren, Sun, Pax, Su, Shiqian, Tao, Chenxin, Bin Wang, H. Wang, Wang, Haonan, James Z. Wang, Jinfeng Wang, Wang, Jojo, Letian Wang

Published Nov 14, 2025▲ 197 on Hugging FaceCode ★ 8,419arXiv ↗

88%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel15/20reviewers recommend it
lenient 5/5
medium 10/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
MiroThinker establishes interaction scaling as a powerful third dimension that pushes open-source agents toward commercial performance, though its 600-call trajectories remain unproven as genuine error-correcting reasoning rather than frequency-overfitted long-horizon loops.

Abstract

We present MiroThinker v1.0, an open-source research agent designed to advance tool-augmented reasoning and information-seeking capabilities. Unlike previous agents that only scale up model size or context length, MiroThinker explores interaction scaling at the model level, systematically training the model to handle deeper and more frequent agent-environment interactions as a third dimension of performance improvement. Unlike LLM test-time scaling, which operates in isolation and risks degradation with longer reasoning chains, interactive scaling leverages environment feedback and external information acquisition to correct errors and refine trajectories. Through reinforcement learning, the model achieves efficient interaction scaling: with a 256K context window, it can perform up to 600 tool calls per task, enabling sustained multi-turn reasoning and complex real-world research workflows. Across four representative benchmarks-GAIA, HLE, BrowseComp, and BrowseComp-ZH-the 72B variant achieves up to 81.9%, 37.7%, 47.1%, and 55.6% accuracy respectively, surpassing previous open-source agents and approaching commercial counterparts such as GPT-5-high. Our analysis reveals that MiroThinker benefits from interactive scaling consistently: research performance improves predictably as the model engages in deeper and more frequent agent-environment interactions, demonstrating that interaction depth exhibits scaling behaviors analogous to model size and context length. These findings establish interaction scaling as a third critical dimension for building next-generation open research agents, complementing model capacity and context windows.