Good Papers

Showing Test-time compute Show all papers

71%Highly rated
?Highly ratedVote to see the score

Towards In-Parameter Memory Augmentation for Large Language Models

This survey organizes in-parameter memory augmentation for LLMs by parameter placement and acquisition time to enable reusable parametric knowledge at deployment.

Haoyu Huang, Zhongwei Xie, Jiaxin Bai, Yisen Gao and 5 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 5/5
medium 1/10
strict 0/5
89%Must read
?Must readVote to see the score

Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

Self-generated feedback in long-horizon test-time training causes weight updates that improve synthetic text but degrade real-text prediction, and settlement on independent evidence prevents this failure.

Cheng Luo, Bing Li, Bernard Ghanem

Published Oct 4, 2026 · ▲ 19 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5
88%Must read
?Must readVote to see the score

ALoDLM: Adaptively Looped Diffusion Language Models

ALoDLM applies token-adaptive latent recurrence to diffusion language models, allocating computation by difficulty to close the quality gap with autoregressive models at 1.7B and 8B scales.

Liancheng Fang, Zhuowei Li, Youngeun Kim, Tianchen Zhao and 9 more

Published Oct 3, 2026 · ▲ 55 on Hugging Face · Code ★ 4

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

ReSolve: Reusing Candidate Reasoning through Selective Generative Moderation

ReSolve reuses candidate reasoning via selective generative moderation to boost math accuracy and cut token use versus voting and self-consistency.

Bangji Yang, Jiajun Fan, MA Hongba, Xi Zhu and 5 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Learning to Predict Distributions over Weight Updates for Test-Time Adaptation

Query-conditioned hypernetworks predict distributions over LoRA weight updates from input queries, enabling test-time scaling via sampled adapted models that outperform deterministic and token-sampling baselines.

Azal Ahmad Khan, Keshav Ramji, Tahira Naseem, Ali Anwar and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
80%Must read
?Must readVote to see the score

LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

AutoTTS automatically discovers test-time scaling strategies via environment-driven controller synthesis, improving LLM reasoning accuracy-cost tradeoffs over manual baselines at minimal cost.

Tong Zheng, Haolin Liu, Chengsong Huang, Huiwen Bao and 9 more

Published May 8, 2026 · 0 citations · ▲ 70 on Hugging Face · Code ★ 176

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

TTCS: Test-Time Curriculum Synthesis for Self-Evolving

TTCS uses a co-evolving question synthesizer and reasoning solver to build test-time curricula that stabilize self-updates and improve reasoning on hard math benchmarks.

Chengyi Yang, Zhiyi Xiang, Yunbo Tang, Zongpei Teng and 4 more

Published Jan 30, 2026 · 0 citations · ▲ 35 on Hugging Face · Code ★ 53

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

When One LLM Drools, Multi-LLM Collaboration Rules

Multi-LLM collaboration outperforms single LLM reasoning on tasks where individual models fail, demonstrating collective rule over solo drooling.

Shangbin Feng, Wenxuan Ding, Alisa Liu, Zifeng Wang and 9 more

Published 2026 · 1 citation

– ReadersNo votes yet. 1 from authors or colleagues not counted
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Efficient Test-Time Scaling via Self-Calibration

Self-Calibration distills self-consistency confidence into LLMs for reliable single-pass estimation, enabling confidence-based early stopping that improves MathQA accuracy to 83.6 with 16 samples.

Chengsong Huang, Langlin Huang, Leng, Jixuan, Jiacheng Liu and 1 more

Published Feb 25, 2025 · 1 citation · ▲ 15 on Hugging Face · Code ★ 22

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Mean-Field Parallel Decoding for Discrete Diffusion Language Models

Tamim Zoabi, Ameen A Ali, Liran Ringel, Lior Wolf

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production

Alexandre Galashov, Matt Jones, Nan Rosemary Ke, Yuan Cao and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Early Signals, Strong Decisions: Prefix-Guided Sampling for Parallel Test-Time Scaling

Jie Ren, Jonathan S Rosenfeld, Neil Thompson

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Test-time Scaling for Diffusion Language Models with Frequency-Aware Remasking

Bowen Zuo, Yue Yu, Dongruo Zhou, Yinglun Zhu

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Relaxation-Aligned State Control for Test-Time Scaling in Generative Combinatorial Optimization

Bohao Li, Ying Li, Pei He, Yangming Guo

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Confidence-Calibrated Inference Expansion for Evaluator-Guided Test-Time Reasoning

Weida Liang, Kenji Kawaguchi

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A Model of Diverse Sampling from Language Models

Manuel Prada-Corral, Yahya Emara, Timothy O'Donnell, Ryan Cotterell and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Training Quality Determines Efficiency Boundaries in Test-Time Reasoning

MD Azizul Hakim

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Test-Time Prompt-Agnostic Decomposition

Junze Wang, Lei Fan, Dezheng Zhang, Donglin Di and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

When Does Structure Help? Statistical Tradeoffs for Structured Reverse Processes in Diffusion Large Language Models

Ruofeng Yang, Jingyuan Liu, Shuai Li

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
Show 20 more papers