Good Papers

Showing papers from University of Wisconsin-Madison Show all papers

93%Must read
?Must readVote to see the score

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

RL post-training yields progress advantage, a log-ratio that recovers optimal step-level advantage without dedicated reward models, outperforming trained alternatives across agent benchmarks.

Changdae Oh, Wendi Li, Seongheon Park, Samuel (Min-Hsuan) Yeh and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 13 on Hugging Face · Code ★ 12

100% Readers1 of 1 upvoted
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

ECHO: Terminal Agents Learn World Models for Free

ECHO trains terminal agents to predict environment responses for dense supervision, doubling GRPO pass@1 on TerminalBench-2.0.

Vaishnavi Shrivastava, Ahmed Awadallah, Dimitris Papailiopoulos

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
45%Niche pick
?Niche pickVote to see the score

Intra-Option Fitted Q-Evaluation: Evaluating Hierarchical Policies from Non-Hierarchical Data

Yunfu Deng, Josiah Hanna

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

CHORD: Cross-Model Hallucination Detection via Relational Graph Discrimination

Yongxin Deng, Zhen Fang, Guansong Pang, Sharon Li and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling

Nicholas Corrado, Wenyuan Huang, Josiah Hanna

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Spectral-Spatial Interpretation

Haotian Ma, Ruqi Yang, Philip Townsend

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Sample Complexity of Linear Regression under Random-Location Coordinate Corruptions

Ilias Diakonikolas, Jingyi Gao, Daniel Kane, Thanasis Pittas

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

(Strongly) Replicable Distribution Testers imply High Probability Distribution Testers

Ilias Diakonikolas, Jingyi Gao, Daniel Kane, Sihan Liu and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Test-Time Sequential Steering of Diffusion Models via Preconditioned Crank-Nicolson

Joel Keller, Taos Transue, Qin Li, Shih-Hsin Wang and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents

Yun-Shiuan Chuang, Ruixuan Tu, Chengtao Dai, You Li and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Is the Importance Ratio Necessary for Stable Reinforcement Learning in LLMs?

Shuibai Zhang, Junhyuck Kim, Gyeongman Kim, Jaewoong Cho

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Entropy Dynamics of Agent Reinforcement Learning

Wendi Li, Shawn Im, Sharon Li

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Measuring and Strengthening Behavioral Suppression in Language Models

Luxi (Lucy) He, Pengcheng Jiang, Jifan Zhang, Jiawei Han and 1 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Fast Sandwich Products in Clifford Algebra

Travis Pence, Daisuke Yamada, Jiaqi Mo, Chanyoung Moon and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

NPUsper: Eliminating Redundant Computation for Real-Time Whisper on Mobile NPUs

NPUsper eliminates redundant Whisper computation on mobile NPUs via online hallucination detection and chunked decoding to cut latency, TTFT, and power.

Hojeong Lee, Si H Lee, Sungwon Woo, Chengpo Yan and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
80%Must read
?Must readVote to see the score

Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

RePlaid, a continuous diffusion language model aligned with modern discrete architectures, achieves scaling laws rivaling discrete diffusion and sets a continuous diffusion perplexity record of 22.1 on OpenWebText.

Zhihan Yang, Wei Guo, Shuibai Zhang, Subham Sahoo and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 2/5
70%Highly rated
?Highly ratedVote to see the score

Adaptive Delayed-Update Cyclic Algorithm for Variational Inequalities

ADUCA is a parameter-free cyclic algorithm for Minty variational inequalities that uses delayed operator updates to avoid line searches and achieves near-optimal global oracle complexity.

Yi Wei, Xufeng Cai, Jelena Diakonikolas

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 5 of 20 reviewers recommend it
lenient 1/5
medium 3/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Online Localized Conformal Prediction

OLCP combines online adaptation with covariate localization for efficient online conformal prediction under heterogeneity, with OLCP-Hedge selecting bandwidths via online expert aggregation; both achieve valid long-run coverage with narrower prediction sets.

Yuheng Lai, Garvesh Raskutti

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
88%Must read

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

AgentGrad improves multi-agent prompt optimization via sequential intervention and semantic gradient clustering, achieving state-of-the-art performance with 2.5x faster optimization and 21.8% lower cost.

Jaewon Chu, jinwoo seo, Jaewon Cho, Jeehye Na and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 107 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

Generating the Unheard: Phylogeny-Guided Latent Generation for Ancestral Sound Reconstruction

This framework generates ancestral bird vocalizations by inferring decodable VAE latents guided by phylogenetic traits, achieving genuine generation and naturalistic audio quality.

Tianyi Xu, Shrinaath Narasimhan, Evan Gorstein, Santiago Perea and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers

Task-vector geometry governs dual inference: in-distribution tasks use convex combinations of learned vectors, while out-of-distribution tasks use nearly orthogonal extrapolative subspaces.

Hao Yan, Haolin Yang, Yiqiao Zhong

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

Testable Learning of General Halfspaces under Massart Noise

A testable learning algorithm learns general Massart halfspaces under Gaussian marginals with quasi-polynomial complexity matching SQ lower bounds.

Ilias Diakonikolas, Giannis Iakovidis, Daniel Kane, Sihan Liu

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 2/5
medium 4/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Polynomial-Time Robust Multiclass Linear Classification under Gaussian Marginals

Multiclass linear classification under Gaussian marginals achieves polynomial-time robust learning via pairwise and localization frameworks, yielding near-optimal error bounds and exposing perceptron limitations.

Ilias Diakonikolas, Giannis Iakovidis, Mingchen Ma

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 2/5
medium 6/10
strict 3/5
89%Must read
?Must readVote to see the score

Combating Data Laundering in LLM Training

Data laundering transforms proprietary data to hide LLM training traces, and Synthesis Data Reversion restores detection by synthesizing likely transformed queries via goal-detail abstraction.

Muxing Li, Zesheng Ye, Sharon Li, Feng Liu

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
83%Must read
?Must readVote to see the score

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

Hide-and-Seek formulates VLA failure detection as coarsely supervised learning to localize failure-indicative actions from trajectory-level labels alone via contrastive objectives, achieving state-of-the-art multi-task detection with practical accuracy-timeliness trade-offs.

Seongheon Park, Wendi Li, Changdae Oh, Samuel (Min-Hsuan) Yeh and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
89%Must read
?Must readVote to see the score

Multi-Head Recurrent Memory Agents

Multi-Head Recurrent Memory partitions recurrent agent memory into independent heads to prevent overwriting, boosting long-context retention from under 30% to 74% at 896K tokens.

Jiatong Li, Samuel (Min-Hsuan) Yeh, Sharon Li

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 2/5
80%Must read
?Must readVote to see the score

Tracing Agentic Failure from the Flow of Success

OAT trains neural controlled differential equations on successful agent trajectories to detect failure steps without failure annotations, outperforming prompting baselines by up to 20% F1 with 200-5000x speedup.

Samuel (Min-Hsuan) Yeh, Yiwen Zhu, Shaleen Deep, Sharon Li

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 16 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
80%Must read
?Must readVote to see the score

Corrective Diffusion Language Models

Standard diffusion language models lack reliable token correction, so a correction-oriented post-training principle improves iterative refinement and outperforms masked diffusion baselines.

Shuibai Zhang, Fred Peng, Yiheng Zhang, Jin Pan and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 17

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Finding Koopman Invariant Subspaces via Personalized PageRank

Personalized PageRank detects Koopman-invariant dictionary subspaces via EDMD zero blocks with finite-sample guarantees and controls multi-step leakage without assuming invariance.

Hyukpyo Hong, Qin Li, Matthew J Colbrook, Hanbaek Lyu

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 2/5
medium 8/10
strict 1/5
72%Highly rated
?Highly ratedVote to see the score

Statistical Query Lower Bounds for Smoothed Agnostic Learning

Smoothed agnostic halfspace learning requires SQ complexity d^Ω(1/σ²+log(1/ε)), nearly matching the best known upper bound.

Ilias Diakonikolas, Daniel Kane

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 1/5
medium 4/10
strict 3/5
83%Must read
?Must readVote to see the score

Scalable Derivative Gaussian Processes via Exact Gradient Reduction

TERA introduces exact gradient reduction for derivative GPs, reducing inference cost to O(dm²+m⁶) per target with flat scaling in dimension d while preserving the model and improving predictive accuracy.

Hyunseok Seung, Matthias Katzfuss

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 2/5