Good Papers

Showing Mathematical reasoning Show all papers

83%Must read
?Must readVote to see the score

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

A new benchmark shows LLMs lack structural mathematical understanding with discovery as the key bottleneck, and a primitive-guided self-distillation framework repairs reasoning to boost performance.

Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin and 6 more

Published Oct 1, 2026 · 0 citations · ▲ 10 on Hugging Face · Code ★ 8

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
45%Niche pick
?Niche pickVote to see the score

Auxiliary Clues Aware’s Geometry Problem Solving

Qigong Lei, Qingsong Wang, Wang Lin, Qianyi Yang and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Cross-Channel Agreement Beats Consensus: Compositional Verification for Geometry Reasoning

ShaoWei Huang, Zefei Gao, Yuting Yu, Hao Yang and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

From Walls to Synergy: A Joint LLM-Evolution Framework for MILP Solvers

Ziao Guo, Yuan Feng, Chenhao Ying, Junchi Yan

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

L$^2$EAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

Po-Nien Kung, Linfeng Song, Dawsen Hwang, Jinsung Yoon and 9 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Can LLMs Reliably Grade Olympiad Proofs? A Controlled Study of Mathematical Verification with LLMs

Azim Ospanov, Zijin Feng, Ding Ding, Chengwu Liu and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 2/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Diagnosing Math-Reasoning Failure Structure with Milestone Oracles

Zhuohan Wang, Haoran Ma, Tianyu Wu, Yuanlin Duan and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Don’t Learn What You Can Compute: Arithmetic Residual Blocks for Exact Arithmetic in Transformers

Eric Gilerson

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

DAPS: Dependency-Aware Premise Selection for LLM Theorem Proving

Zixuan Chen, Wenyuan Jiang, Junling Wang, Mrinmaya Sachan and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning

Supervised fine-tuning harms reasoning by memorizing surface correlations rather than theorem application; Theorem-SFT improves MATH and GeoQA scores by teaching explicit rule invocation.

Ruiying Peng, Mengyu Yang, Jing Lei, Xiao-Hui Li and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 2/5
88%Must read
?Must readVote to see the score

Escaping the Cognitive Well: Efficient Competition Math with Off-the-Shelf Models

An inference pipeline using off-the-shelf models and conjecture extraction with context detachment achieves state-of-the-art IMO-style math performance at much lower cost by escaping the Cognitive Well.

Xingyu Dang, Rohit Agarwal, Rodrigo Porto, Anirudh Goyal and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic

Transformers learn arithmetic skills non-sequentially via correlational matching, causing mixing errors and poor robustness to distribution shifts even when scaled.

Xingyu Zhao, Darsh Sharma, Rheeya Uppaal, Yiqiao Zhong

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
83%Must read
?Must readVote to see the score

COMPOSE: Composing Future Theorems from Citations and Formal Structure

COMPOSE generates future theorem-like claims by combining citation graphs with formal theorem dependencies, outperforming baselines on retrieval and evaluation.

David Busbib, Michael Werman

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5