Good Papers

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Parallel-R1 uses a progressive curriculum combining supervised warmup and reinforcement learning to train large language models in parallel reasoning, achieving significant gains on math benchmarks by treating parallel thinking as a temporary exploration scaffold that unlocks higher final performanc

Zheng, Tong, Hongming Zhang, Wenhao Yu, Xiaoyang Wang, Dai, Runpeng, Rui Liu, Huiwen Bao, Huang, Chengsong, Heng Huang, Dong Yu

Published Sep 9, 2025▲ 105 on Hugging FaceCode ★ 265arXiv ↗

83%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel13/20reviewers recommend it
lenient 4/5
medium 9/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Parallel-R1 uses RL and a progressive curriculum to turn multi-path reasoning into a powerful cold-start scaffold that yields striking AIME25 gains, though critics note it remains untested beyond math, lacks equal-compute baselines, and treats parallel thinking…

Abstract

Parallel thinking has emerged as a novel approach for enhancing the reasoning capabilities of large language models (LLMs) by exploring multiple reasoning paths concurrently. However, activating such capabilities through training remains challenging, as existing methods predominantly rely on supervised fine-tuning (SFT) over synthetic data, which encourages teacher-forced imitation rather than exploration and generalization. Different from them, we propose \textbf{Parallel-R1}, the first reinforcement learning (RL) framework that enables parallel thinking behaviors for complex real-world reasoning tasks. Our framework employs a progressive curriculum that explicitly addresses the cold-start problem in training parallel thinking with RL. We first use SFT on prompt-generated trajectories from easier tasks to instill the parallel thinking ability, then transition to RL to explore and generalize this skill on harder problems. Experiments on various math benchmarks, including MATH, AMC23, and AIME, show that Parallel-R1 successfully instills parallel thinking, leading to 8.4% accuracy improvements over the sequential thinking model trained directly on challenging tasks with RL. Further analysis reveals a clear shift in the model's thinking behavior: at an early stage, it uses parallel thinking as an exploration strategy, while in a later stage, it uses the same capability for multi-perspective verification. Most significantly, we validate parallel thinking as a \textbf{mid-training exploration scaffold}, where this temporary exploratory phase unlocks a higher performance ceiling after RL, yielding a 42.9% improvement over the baseline on AIME25. Our model, data, and code will be open-source at https://github.com/zhengkid/Parallel-R1.