Good Papers

Showing LLM reasoning & chain of thought Show all papers

89%Must read

Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

On-policy power distillation trains models to generate sharpened answers directly, improving single-sample math reasoning by up to 27.3 points and outperforming multi-candidate sampling and reward-based methods.

Erfan Baghaei Potraghloo, Seyedarmin Azizi, Arya Fayyazi, Saeid Shokoufa and 3 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5
88%Must read
?Must readVote to see the score

Base Models Can Reason By Taking a Cue From Training Data

Fixing initial token cues in base models boosts reasoning to match RL performance, with effects traced to training data associations that can be causally edited.

Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat and 2 more

Published Oct 5, 2026 · ▲ 12 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
83%Must read
?Must readVote to see the score

What Matters for Latent Reasoning with Flow Matching

FLaRe uses flow matching for latent reasoning that is useful, diverse, explainable, refinable and efficient, reaching 97% of explicit chain-of-thought accuracy at 25% latency.

Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

Published Oct 5, 2026 · ▲ 10 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 0/5
80%Must read
?Must readVote to see the score

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Efficient reasoning training differs in impact: faithfulness usually drops due to inconsistency, but monitorability remains robust.

Samuel Lewis-Lim, Xingwei Tan, Mario Sänger, Zhixue Zhao and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 13 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
88%Must read
?Must readVote to see the score

Language Models that Play Chess and Explain Their Moves

Queen, a 4B-parameter chess-language model, plays at grandmaster level and explains moves via cross-attention to a silent expert encoder and iterative Bellman-style explanation distillation, surpassing larger frontier models.

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Published Oct 2, 2026 · 0 citations · ▲ 31 on Hugging Face · Code ★ 18

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Hierarchical Continuous Diffusion Language Models

HC-DLM couples discrete token generation with a continuous latent trajectory via a unified variational denoising objective, outperforming diffusion baselines on Sudoku, Countdown, and language modeling.

Hui Ren, Zihan Li, Chang Liu, Huidong Liu and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 89 on Hugging Face · Code ★ 57

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
80%Must read
?Must readVote to see the score

Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

BDA treats multi-LLM council moves as observations of a per-agent reliability model to yield calibrated answer probabilities and robustly handle persistent adversaries without extra LLM calls.

Ionel Eduard Stan, Paolo Napoletano

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
91%Must read

Does Learning Protein Folding Generalize to Broader Reasoning?

Post-training on protein-folding data via discrete answers and continuous geometry improves structure prediction and broad reasoning across ten benchmarks.

Yong Liu, Zhanpeng Shi, Yizhou Dang, Zhongyue Zhang and 3 more

Published Sep 30, 2026 · 0 citations · ▲ 120 on Hugging Face · Code ★ 28

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
86%Must read
?Must readVote to see the score

Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation

Interpolated Policy Distillation mixes student and teacher token distributions to balance trajectory quality and learnability, outperforming off-policy and on-policy distillation across reasoning benchmarks.

Youxu Shi, Yifan Sun, Dacheng Yin, Haomiao Tang and 4 more

Published Sep 29, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
86%Must read
?Must readVote to see the score

TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment

TeacherGRPO aligns teachers to student distributions via reinforcement learning to overcome reasoning distillation's Gap Curse and improves student performance.

Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng and 4 more

Published Aug 20, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
72%Highly rated
?Highly ratedVote to see the score

On the Geometry of On-Policy Distillation

On-policy distillation updates occupy a sparse, low-dimensional parameter subspace that is functionally sufficient and geometrically distinct from supervised fine-tuning and reinforcement learning.

Zhennan Shen, Yanshu Li, Qingyu Yin, Chak Tou Leong and 5 more

Published Jun 5, 2026 · 0 citations · ▲ 75 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 1/5
80%Must read
?Must readVote to see the score

RelayLLM: Efficient Reasoning via Collaborative Decoding

RelayLLM enables small language models to dynamically invoke large models for critical reasoning tokens via collaborative decoding, reducing costs by 98.2% while achieving 49.52% accuracy.

Chengsong Huang, Tong Zheng, Langlin Huang, Jinyuan Li and 2 more

Published Jan 8, 2026 · 0 citations · ▲ 30 on Hugging Face · Code ★ 41

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

DeepSeek-V3.2 improves efficiency via sparse attention, scaled reinforcement learning matching GPT-5, and agentic synthesis, with a special variant surpassing GPT-5 and reaching gold-medal IMO and IOI levels.

DeepSeek-AI, Aixin Liu, Mei, Aoxue, Lin, Bangcai and 36 more

Published Dec 2, 2025 · 8 citations · ▲ 274 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs

SoftCoT uses a fixed assistant and projection module to generate soft reasoning tokens that boost LLM reasoning via parameter-efficient fine-tuning.

Yaogeng Xu, Xu Guo, Zhiwei Zeng, Chunyan Miao

Published 2025 · 14 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Continuous Audio Thinking for Large Audio Language Models

Gyojin Han, DongJae Lee, Changho Choi, Jongsuk Kim and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Hypothesis generation and updating in large language models

Huadong Xiong

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Reasoning with Sampling: Cutting at Decision Points

Felix Zhou, Anay Mehrotra, Quanquan C Liu

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

A Solvable Model of Chain-of-Thought in In-Context Learning

Kaito Takanami, Cengiz Pehlevan

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Contractive Restoring Flows: Robust Reasoning Distillation via Orbital Stability

Dongqi Zuo, Yuanyuan Wang, Chuan Zhou, Haoxuan Li and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Programmatic Reasoning with Structural Schema: A Unified Framework for Multi-Table Inference

Jialin Chen, Brandon Mayer, Michael Galkin, Sami Abu-El-Haija and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
Show 20 more papers