Good Papers

Showing papers from Amazon Show all papers

78%Highly rated
?Highly ratedVote to see the score

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Tuna-2 replaces vision encoders with patch embeddings for end-to-end pixel-space multimodal understanding and generation, achieving state-of-the-art results that outperform encoder-based designs at scale.

Zhiheng Liu, Weiming Ren, Xiaoke Huang, Shoufa Chen and 11 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 70 on Hugging Face · Code ★ 756

– ReadersNo votes yet. 1 from authors or colleagues not counted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
67%Highly rated
?Highly ratedVote to see the score

Incentivizing Agentic Retrieval for Disease-Centric Clinical Case Search via Trajectory Memory

Jie Lin, Xiang Liu, Lihao Liu, Liansheng Wang

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Base Items Overfit, New Items Underfit: Hidden Cost of Joint Training in Incremental Adaptation

Rishabh Agrawal

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
67%Highly rated
?Highly ratedVote to see the score

AndroidReality: How Far Are Mobile Agents from the Real World?

Xiaoou Liu, Longchao Da, Hanyang Chen, Yuan Ling and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SPLICE: Structured Prompt Local Iterative Combinatorial Evolution

Dr. Anish Acharya, Phillip Studans, Amit Dhanda, Ninad V Rao and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems

Tunyu Zhang, Zihao Zhao, Yusong Zhao, Haizhou Shi and 4 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Exact Unlearning via Quantized Sufficient Statistics

Ami Tavory, Shripad Gade, Tal Sarig, Noam Touitou and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

From Click Imitation to Transition Equivalence: Rethinking Supervision for GUI Agents

Zhiming Lin, Tianxiang Xu, zizhao zhang, Yixue Liu and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

Beyond IPS: Reliable Counterfactual Evaluation in Multi-Stage Ad Systems without Logged Propensities

Mohsen Malmir, Mohamed A Radwan, houssam nassif, Murat Bayir

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

BayesJudge: Uncertainty-Aware Bayesian Meta-Evaluation of Human and LLM Judgments

Jiahao Zhang, Pengbin Feng, Chunlei Meng, Hang He and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Scaling Arbitrary Architectures and Optimizers with Automatic Parameterization

Shikai Qiu, Charlie Chen, Andres Potapczynski, Martin Marek and 1 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Embedded-Arena: Building Hardware-in-the-Loop Coding Agents to Run AI on Microcontrollers

Zhihan Zhang, Alexander Le Metzger, Jiuyang Lyu, Chun-Cheng Chang and 9 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Margin Dynamics for Large Language Model Alignment

Xingzi Xu, Saygin Seyfioglu, Karim Bouyarmane

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CompJudge: Fine-Grained Comparative Evaluation using Multimodal LLM for Subject-Driven Generation

Nam Hyeon-Woo, Wenbin Ouyang, Ciprian A Corneanu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

ANCHOR: Audio-Visually Grounded Chain-of-Thought Reasoning Benchmark

Joel Julin, Souraja Kundu, Liza Dahiya, George Z Wei and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

MaRiO: Multi-agent Collaborative Reasoning via Shared Observations in MLLMs

Nathaniel Redmond, Fidel Omar Tito Cruz, Devansh Sharma, Shehreen Azad and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Reward Shaping to Improve Language Model Query Generation

Shicheng Liu, Zeyu Zhang, LEI LU, Kexuan Sun and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

fev-bench: A Realistic Benchmark for Time Series Forecasting

Oleksandr Shchur, Abdul Fatir Ansari, Ali Caner Turkmen, Lorenzo Stella and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · Code ★ 179

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CREF: Forecasting Benchmarks for the Age of Agents

Andreas Auer, Abdul Fatir Ansari, Oleksandr Shchur, Xiyuan Zhang and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Beyond Task Success: Probing Cognitive Primitives in Web Agents

Xunjian Yin, Tianchen Guan, Jinao Wang, Weili Cao and 7 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Reasoning-based Spatial Prior (RSP): Learning Spatial Priors from Multimodal LLMs for Object Detection

Cagri Gungor, Qingshuang Chen, Hongda Mao, Chi Zhang and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

V-LUMEN: Visual Lookup Memory for Embedding Scaling in Vision-Language Models

Miso Choi, Daekeun Kim, Eunji Kim, Jungbeom Lee

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Koopman Generative Operators for Efficient Probabilistic Time-Series Forecasting

Raz Marshanski, Liran Nochumsohn, Mayank Jauhari Iitr, Boris Oreshkin and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Surjective Pseudo-Invertible Neural Networks

Yamit Ehrlich, Amit Arad, Nimrod Berman, Assaf Shocher

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Beyond Copy-Paste: How Well Do Subject-Driven Video Models Understand Their Subjects?

Zun Wang, Kenan Deng, Daniel Blackburn, Linlin Lu and 4 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ContractBench: Can LLM Agents Preserve Observation Contracts?

Jicheng Wang, Yifeng He, Zili Wang, Hanwen Xing and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DynaPFN: Zero-Shot Dynamical System Forecasting with Tabular Prior-Fitted Networks

Chiara Roverato, Joseph Cotnareanu, Pablo Piantanida, Boris Oreshkin and 1 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SLiDE: Structured Linear Dynamics for Forecasting with Exogenous Inputs

Sebastian Pütz, Theodore Glavas, Benjamin Schäfer, Boris Oreshkin and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Fast-dLLM++: Fr\'{e}chet Profile Decoding for Faster Diffusion LLM Inference

Siva Rajesh Kasa, Yasong Dai, Sumit Negi, Hongdong Li

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Steering Vectors as a Training Signal in LLM Post-Training

Tiejin Chen, Maunil R Vyas, Huaiyuan Yao, Hua Wei

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Whole-Body Compliant Control via Learned Force-Regulation Modules

Diego Aldarondo, Aadhithya Iyer, Daniel Giebisch, Nina Mortensen and 3 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

nnTrace: Detecting and Localizing Silent Bugs in Distributed Training

Haitian Jiang, Shaowei Zhu, Zhen Zhang, Zhenyu Song and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Permute-then-Adapt: Weak-to-Strong Contrastive Image--Text Adaptation

Jinhao Li, Sarah Erfani, Lei Feng, Guangrui Li and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MID: Mask-Image Distributional Divergence for Evaluating Medical Image Segmentation

Vincenzo Marcianò, XIAOMING ZHANG, Gianluca Guglielmo, Sebastien Ourselin and 2 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Alleviating Hallucination with Training-Free Uncertainty-Guided Steering

Litian Liu, Yubing Jian, Qiqi Hou, Reza Pourreza and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models

Gaurav Patel, Jun Fang, Greg Ver Steeg, Qiang Qiu and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Firefly: Illuminating Verified Real-world Tool Call Data Generation

Yuxuan Lu, Ziyi Wang, Yingzhou Lu, Yisi Sang and 11 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

A$^2$IQL: Adaptive Asymmetric Implicit Q Learning for Automated Warehouse Consolidation

Guangyi Liu, Andrea Angiuli, Mirko Ristivojevic, Joseph W Durham and 2 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

ALIGN-Rec: Continual Recommendation under Heterogeneous Unlearning Requests

Nitin Bisht, Sumit Bisht, Tong Zhang, Yu Yang and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Network of Theseus (Like the ship)

Network of Theseus progressively replaces guide network modules with a different target architecture via representational alignment, preserving performance across vastly different deployed architectures.

Vighnesh Subramaniam, Colin Conwell, Boris Katz, Andrei Barbu and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Stay Fair! Ensuring Group Fairness in Diffusion Models Across Guidance Scales

StayFair decomposes diffusion bias into model and guidance components, deriving a guidance-scale-invariant fairness condition and algorithms that maintain group fairness across all guidance scales without quality loss.

Myeongsoo Kim, Eunji Kim, Minwoo Chae, Sangwoo Mo

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation

MURPHY extends GRPO to multi-turn code generation via feedback-conditioned rollout trees with retrospective credit assignment, achieving up to 6% absolute pass@1 gains over prior methods.

Chanakya Ekbote, Vijay Lingam, Sujay Sanghavi, Luke Huan and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent

AgentArk distills multi-agent reasoning into a single LLM via hierarchical strategies, achieving multi-agent performance with single-agent efficiency.

Yinyi Luo, Yiqiao Jin, Weichen Yu, Mengqi Zhang and 5 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 8 on Hugging Face · Code ★ 216

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Mean Testing under Truncation beyond Gaussian

Under truncation hiding an ε-fraction of mass, mean testing faces a bias floor of order ν ε^{1−1/p}; above it a second-order test achieves near-optimal sample complexity, while median regularity restores classical √d testing rates.

Yuhao Wang, Roberto I Oliveira, Themis Gouleakis

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection

CompactAttention accelerates chunked prefill via block-union KV selection that builds minimal per-group block tables to eliminate KV copy overhead, achieving up to 2.72x attention speedup with near-dense accuracy on 128K contexts.

Jiwon Song, Dongwon Jo, Beomseok Kang, jae-joon kim

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 12 on Hugging Face · Code ★ 7

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization

Epiplexity guides data selection and synthetic generation to improve out-of-distribution transfer by favoring structurally rich training data.

Ellen Su, Andres Potapczynski, Shikai Qiu, Edward Hughes and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
89%Must read
?Must readVote to see the score

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents

OpenWebRL enables open online RL for visual web agents, with a 4B model reaching 67% Online-Mind2Web and 64% DeepShop success using minimal initialization data.

Rui Yang, Qianhui Wu, Yuxi Chen, Hao Bai and 6 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 19 on Hugging Face · Code ★ 52

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

TwinRouterBench introduces static and live dynamic tracks to benchmark LLM routing at agent step-level using deterministic scoring and live execution on SWE-bench.

Pei Yang, Wanyi Chen, Tongyun Yang, Pengbin Feng and 11 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

ExComm: Exploration-Stage Communication for Error-Resilient Agentic Test-Time Scaling

ExComm detects cross-agent factual conflicts during exploration to correct errors via soft belief updates and trajectory diversification, improving test-time scaling by up to 5.7%.

Woomin Song, Beomjun Kim, Daewon Choi, Sai Muralidhar Jayanthi and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 0/5
86%Must read
?Must readVote to see the score

Consilience for Verifier-Free Test-Time Scaling

Confidence-based verifier-free test-time scaling fails on complex tasks because high initial confidence signals no exploration; consilience selects rollouts by requiring low early but high final confidence, improving reasoning and coding.

Lecheng Kong, Like Hui, Haitao Mao, Luke Huan

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
88%Must read
?Must readVote to see the score

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering

Visual Sparse Steering trains sparse autoencoders on frozen CLIP activations to build label-free steering vectors that improve zero-shot classification by up to 4.12 percent via centroid-deviation steering with reconstruction-error gating.

Gerasimos Chatzoudis, Zhuowei Li, Gemma Moran, Hao Wang and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

92%Must read
?Must readVote to see the score

Computer Use at the Edge of the Statistical Precipice

A 1MB replay script outperforms frontier agents on static benchmarks because of flawed environment design and evaluation; the paper proposes PRISM principles, DigiWorld, and hierarchical bootstrap aggregation to fix both.

Pierluca D Oro, Sneha Silwal, William R Wong, Yuxuan Sun and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs

AstraFlow is a dataflow-oriented RL system for agentic LLMs that decouples rollout, dataflow, and training to enable multi-policy collaborative training with 2.7x faster training.

Haizhong Zheng, Yizhuo Di, Jiahui Wang, Shuowei Jin and 6 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 15 on Hugging Face · Code ★ 105

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning

SWAP assigns step-level length penalties based on reasoning contribution, cutting chain-of-thought length by 64.3% and boosting accuracy 5.7% over base models.

Xintong Li, Sha Li, Rongmei Lin, Hongye Jin and 9 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-Training

Pass@k optimization degrades pass@1 via prompt interference, as its gradients conflict by upweighting negatively interfering, low-success prompts.

Anas Barakat, Souradip Chakraborty, Khushbu Pahwa, Amrit Singh Bedi

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

MARS: Harmonizing Multimodal Convergence via Adaptive Rank Search

MARS introduces adaptive rank search balancing multimodal convergence dynamics via dual scaling laws to optimize low-rank fine-tuning of multimodal large language models.

Minkyoung Cho, Insu Jang, Shuowei Jin, Zesen Zhao and 4 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Align-RAG: Alignment Is All You Need for TSFM In-Context Learning

Align-RAG applies closed-form amplitude rescaling and phase shifts to retrieved windows, outperforming trained adapters across frozen time-series foundation models without training.

Mohammad Asadi, Soheil Hor, Bardiya Akhbari, Jack W O'Sullivan and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Trust, but Don’t Verify: Epistemic Blind Spots in LLM Source Evaluation

LLMs detect fabricated statistics in isolation but ignore numeric validity during multi-source synthesis, weighing sources by analytical register rather than accuracy.

Rohan N Pradhan, Steve Goley

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use

CM2 replaces verifiable outcome rewards with checklist rewards for multi-turn tool-use RL, improving 8B models by 8, 12 points on agent benchmarks using simulated environments.

Zhen Zhang, Kaiqiang Song, Sean Wang, Yebowen Hu and 10 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph

This paper formalizes test-time scaling as an optimizable multi-LLM collaboration graph and proposes Agent-REINFORCE to search compute-optimal architectures under budget constraints. Experiments show it outperforms baselines in efficiency and finds graphs balancing accuracy with latency.

Fali Wang, Jihai Chen, Shuhua Yang, Runxue Bao and 6 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
91%Must read
?Must readVote to see the score

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

PACE uses two-timescale self-evolution to let frozen small language models improve agents via validated prompt and control updates, outperforming baselines on 12 settings by up to 9.2%.

Chen Ling, Pei Chen, Xiangchen Guan, Jiaming Qu and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
83%Must read
?Must readVote to see the score

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

AME-TS guides sparse mixture-of-experts routing via temporal structure descriptors to improve forecasting accuracy and specialization stability with fewer activated parameters.

Rui Wang, Renhao Xue, Ray Razi, Huan Song and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory

Lean4Agent uses Lean4 to formally model and verify agent workflows, with verified workflows outperforming failing ones by 11.94% and LeanEvolve improving SWE performance by 7.47%.

Ruida Wang, Jerry Huang, Pengcheng Wang, Xuanqing Liu and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 7 on Hugging Face · Code ★ 31

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

LACE: Lattice Attention for Cross-thread Exploration

LACE enables parallel reasoning paths to share insights and correct errors during inference via cross-thread attention, improving reasoning accuracy by over 7 points.

Yang LI, Zirui Zhang, Yang Liu, Chengzhi Mao

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Hyper Hawkes Processes: Interpretable Models of Marked Temporal Point Processes

Hyper Hawkes processes extend classical Hawkes models via latent spaces and hypernetworks for flexible, interpretable marked temporal point process predictions.

Alex Boyd, andrew warrington, Taha Kass-Hout, Parminder Bhatia and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Learning How to Cube

A neuro-symbolic framework trains a 4B-parameter model via MCTS-curated preference data and two-stage post-training to learn SAT cubing heuristics matching top symbolic methods.

Ferhat Erata, Sam Kouteili, Thanos Typaldos, Timos Antonopoulos and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

SPANUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

SPANUQ is a lightweight probe that estimates span-level LLM generation uncertainty via hidden-state distillation, outperforming sampling methods with 10, 20x speedups and 0.910 F1 span detection.

Yimeng Zhang, Yingying Zhuang, Ziyi Wang, Yuxuan Lu and 11 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
80%Must read
?Must readVote to see the score

CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents

CORVUS decouples file reads from observations via synchronized registries, cutting input tokens by 9-50% and reasoning cycles by up to 37% while preserving pass rates.

Mingwei Zheng, David OBrien, Siwei Cui, Pardis Pashakhanloo and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

SkillsBench benchmarks agent skills across 87 tasks, finding curated skills boost pass rates by 16.6 points, with focused small bundles often outperforming larger ones.

Xiangyi Li, Yimin Liu, Wenbo Chen, Shenghan Zheng and 36 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Solaris: Building a Multiplayer Video World Model in Minecraft

Solaris introduces a multiplayer video world model and automated data system for Minecraft, collecting 12.64M frames to enable consistent multi-agent simulation via staged training that outperforms single-player baselines.

Oscar Michel, Georgy Savva, Daohan Lu, Suppakit Waiwitlikhit and 6 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Online Data Selection for Instruction Tuning via Gaussian Processes

GAIA casts instruction-tuning data selection as global Gaussian-process utility estimation with adaptive strategy fusion, yielding dynamic-regret guarantees and outperforming batch-constrained baselines.

Jun Wang, Quoc Phong Nguyen, Julien Monteil, Vu Nguyen

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
89%Must read
?Must readVote to see the score

Multi-site PPG: An In-the-Wild Physiological Dataset from Emerging Multi-Site Wearables

Multi-site PPG is an in-the-wild dataset of 350+ hours from earring, ring, watch, and necklace wearables, showing heart-rate errors vary substantially by body site.

Jiayi Shao, Jiaying Ye, ShengYao Liu, Zachary Englhardt and 3 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5
71%Highly rated
?Highly ratedVote to see the score

Model-based Bootstrap of Controlled Markov Chains

A model-based bootstrap for controlled Markov chain transitions yields consistent estimators and valid confidence intervals for offline policy evaluation and optimal recovery.

Ziwei Su, Imon Banerjee, Diego Klabjan

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 0/5
80%Must read
?Must readVote to see the score

Few-Step Diffusion Language Models via Trajectory Self-Distillation

Trajectory self-distillation trains few-step diffusion language models to match full-step trajectories, mitigating factorization error to enable fast, high-quality parallel decoding.

Tunyu Zhang, Xinxi Zhang, Ligong Han, Haizhou Shi and 9 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 9 on Hugging Face · Code ★ 27

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
86%Must read
?Must readVote to see the score

Bandits via Additive Quantized Representations

Residual Quantization maps contexts to discrete additive codes enabling nonlinear contextual bandits with strictly bounded memory, beating linear variants on 11 of 13 datasets and matching heavy retrained baselines with up to 1000x less memory.

Ami Tavory, Noam Touitou, Tal Sarig, Frank Cheng and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
80%Must read
?Must readVote to see the score

An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration

Human soft-labels improve calibration and training stability by regularizing models and mirroring human uncertainty, mainly via regularization rather than correcting mislabeled data.

Maja Pavlovic, Silviu Paun, Massimo Poesio

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

CommunityKV formulates sparse attention as graph community detection to retrieve coherent token groups via constant-time updates, achieving up to 1.71x long-context decoding throughput.

Joe McKenna, Anastasios Alexandridis, Nathan Susanj, Jing Liu

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
70%Highly rated
?Highly ratedVote to see the score

s2n-bignum-bench: A practical benchmark for evaluating low-level code reasoning of LLMs

s2n-bignum-bench evaluates LLM theorem proving on verified industrial cryptographic assembly using HOL Light proof synthesis. It provides a challenging, practically relevant benchmark beyond competition mathematics.

Balaji Rao, Soonho Kong, Juneyoung Lee, Carlo Lipizzi

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 3 on Hugging Face · Code ★ 5

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 5 of 20 reviewers recommend it
lenient 4/5
medium 1/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving

MAPLE trains vision-language-action driving models via latent multi-agent rollout and reinforcement learning, achieving state-of-the-art closed-loop performance without external simulators.

Rajeev Yasarla, Deepti Hegde, Hsin-Pai Cheng, Shizhong Han and 8 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
83%Must read
?Must readVote to see the score

Generative Scenario Rollouts for End-to-End Autonomous Driving

GeRo enables vision-language-action models to generate language-grounded future traffic scenes via autoregressive rollouts, improving Bench2Drive driving scores by 15.7 and success rates by 26.2.

Rajeev Yasarla, Deepti Hegde, Shizhong Han, Hsin-Pai Cheng and 10 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
Show 20 more papers