Good Papers

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning

NPR enables LLMs to self-evolve genuine parallel reasoning via self-distilled reinforcement learning, achieving up to 24.5% accuracy gains, 4.6x speedups, and 100% parallel execution.

Wu, Tong, Liu, Yang, Bai, Jun, Jia, Zixia, Zhang, Shuyi, Lin, Ziyong, Wang, Yanting, Zhu, Song-Chun, Zheng, Zilong

Published Dec 8, 2025▲ 80 on Hugging FaceCode ★ 112arXiv ↗

74%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
NPR delivers a real RL contribution in graph-level branching optimization and substantial speedups, but its claim of native parallel cognition relies on unverified execution flags rather than audited reasoning traces or named benchmarks.

Abstract

We introduce Native Parallel Reasoner (NPR), a teacher-free framework that enables Large Language Models (LLMs) to self-evolve genuine parallel reasoning capabilities. NPR transforms the model from sequential emulation to native parallel cognition through three key innovations: 1) a self-distilled progressive training paradigm that transitions from ``cold-start'' format discovery to strict topological constraints without external supervision; 2) a novel Parallel-Aware Policy Optimization (PAPO) algorithm that optimizes branching policies directly within the execution graph, allowing the model to learn adaptive decomposition via trial and error; and 3) a robust NPR Engine that refactors memory management and flow control of SGLang to enable stable, large-scale parallel RL training. Across eight reasoning benchmarks, NPR trained on Qwen3-4B achieves performance gains of up to 24.5% and inference speedups up to 4.6x. Unlike prior baselines that often fall back to autoregressive decoding, NPR demonstrates 100% genuine parallel execution, establishing a new standard for self-evolving, efficient, and scalable agentic reasoning.