Good Papers

Showing papers from Johns Hopkins University Show all papers

91%Must read

Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification

Automatic taxonomies reveal retrieval and verifier reasoning failures persist across scaled medical retrieve-then-verify systems, showing fundamental open-ended evaluation limits.

Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi and 4 more

Published Sep 24, 2026 · 0 citations · ▲ 12 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
45%Niche pick
?Niche pickVote to see the score

Adaptive Residual Quantization for Memory-Efficient Temporal Action Segmentation

Gerard L Donahue, Guven Gergerli, Ayush Gupta, Reza Ghoddoosian and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Towards Financial World Modeling

Humzah Merchant, Alec Guthrie, Simon Mahns, Randall Balestriero and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation

Biao Xiang, Ali Eshragh, Yuexing Li, Kai Wang

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Multi-Objective Causal Bandits: Minimal Intervention Space and Policy-Level Learning

Muhammad Qasim Elahi, Mahsa Ghasemi, Murat Kocaoglu

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Reading, Not Thinking: Bridging the Modality Gap When Text Becomes Pixels

Kaiser Sun, Xiaochuang Yuan, Hongjun Liu, Chen Zhao and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Contextualized Evaluation of Vision Language Models through Dynamic Interviews

Yijiang Li, Huiqi Zou, Bingyang Wang, Ziang Xiao

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Lost on Campus: Evaluating Embodied Spatial Reasoning of Vision-Language Models in the Wild

Zehan Zheng, Yanyuan Chen, Deming Li, Yutao Tang and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Policy Regret Minimization in Partially Observable Markov Games

Lan Sang, Raman Arora, Thanh Nguyen-Tang

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Can We Trust Item Response Theory for AI Evaluation?

Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Hierarchical Concept Geometry in Language Representations Emerges from Word Co-occurrence

Andres Nava, Matthieu Wyart

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

AnaDiffusion: Anatomically Compositional Latent Diffusion for Controllable 3D Brain MRI Generation

Tracy Han, Lulin Liu, Bangya Liu, Yuanhao Cai and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Context-Aware Autoregressive Image Generation for Emerging Reasoning Properties

Jixuan Ying, Haoyu Liu, Timing Yang, Tingyu Zhu and 6 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Mission Impossible: Diagnosing and Fixing Non-Operative Instruction Following in Image Editing

Guoyizhe Wei, Feng Wang, Alan Yuille, Rama Chellappa

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Transforming Image Editors into Video Editors

Feng Wang, Zijie Li, Ceyuan Yang, Alan Yuille and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

The Narrative Gap: Can LLMs help us Navigate Diverse Narratives Across Languages?

Nikhil Sharma, Kelly Marchisio, Kenton Murray, Ziang Xiao

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

PuppetGait: Generalizing Gait Recognition via 3D Body-Aligned LVM Features

成伟 叶, Dingqiang Ye, Chuanfu Shen, Zirui Zhou and 4 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Focus, Align, and Diffuse: Time-Series-Aware Keyframe Diffusion for Cardiac Dynamic Synthesis from Sparse Observations

Junkai Liu, Haofan Wu, Nay Aung, Joao A Lima and 2 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

\texttt{FEROM}: Frontier Endogenous Reveal-Order Marginal Policy Optimization for Masked Diffusion LMs

Zian Su, Ziyang Huang, Kaiyuan Zhang, Xiangyu Zhang

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Automated Causal Effect Estimation through Self-Evolving AI

Can Wang, Hongyu Zhao, Yiqun Chen

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Characterizing Underrepresentation in Generalizing Causal Survival Estimates

Bolun Liu, Sean McGrath, Yiren Hou, Elizabeth Stuart and 1 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Magnetic Resonance Unpaired Image Translation with Pseudometric Schrödinger Bridges

Shuwen Wei, Samuel Remedios, Zhangxing Bian, Shimeng Wang and 8 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Anatomy-Activated Mixture-of-Experts for 3D Medical Vision-Language Pre-training

Szymon Płotka, Gizem Mert, Pedro R. A. S. Bassi, Wenxuan Li and 8 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Macrocanonical Generator Networks: data-efficient neural surrogates for amortized physics simulation

Niall Jeffrey, Benjamin Wandelt

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Finite-Sample Convergence in Networked Average Reward MARL: Decentralization Pitfalls and Entropy Remedies

Yizhou Zhang, Yashaswini Murthy, Laixi Shi, Adam Wierman

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Decision-Focused Learning in MDPs: An Occupancy Measure Approach

Zihao Zhao, Ashwath K Karunakaram, Ali Eshragh, Yuexing Li and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

NyoomFloat12: Accelerating LLM Inference via Lossless 12-bit Weight Compression

Sylvie Liberman, Xinyu Fang, Tianyi Zhang, Tri Dao and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Neural Compression of Long ADMM Trajectory for Multiparametric Quadratic Program

Liang Wu, Bo Yang, Xu Yang, Honghui Zheng and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Direct Acceleration of Stochastic Root-Finding Without Variance Reduction and Regularization

A dual-anchor mechanism accelerates stochastic root-finding to O(ε⁻³) without variance reduction or regularization, reaching near-optimal O(ε⁻²) for strongly monotone cases.

TaeHo Yoon, Nicolas Loizou

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 2/5
medium 4/10
strict 2/5
80%Must read
?Must readVote to see the score

NEvo: Neural-Guided Evolutionary Video Synthesis for Dynamic Visual Selectivity

NEvo uses neural-guided evolutionary video synthesis to generate brain-region-optimized dynamic stimuli that surpass handcrafted localizers and reveal visual cortex temporal selectivity differences.

Yingtian Tang, Sogand Salehi, Ming Zhou, Amir Zamir and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

How to Interpret Agent Behavior

ACT*ONOMY introduces a three-level taxonomy of 10 actions and 46 subactions for describing autonomous agent behavior at runtime, plus an open repository and automated analysis pipeline that compares behavioral profiles and surfaces failure patterns.

Sophia Gao, Kaiser Sun, Jen-Tse Huang, Katherine Van Koevering and 7 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

PGID: Progressive Guided Inversion and Denoising for Robust Watermark Detection

PGID defends diffusion watermark detectors against removal and forgery attacks by progressively projecting perturbed latents back to their correct regions via guided inversion-denoising cycles, restoring reliable detection without training.

Minh Quoc Duong, Chun Tong Lei, Chun Pong Lau

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Tree Search With Predictions

No algorithm achieves O(log η) search on general trees via distance predictions, but O(k log η) queries work for trees of pathwidth k with optimal complexity.

Michael Dinitz, Bob Dong

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 2/5
80%Must read
?Must readVote to see the score

CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning

CausalSpatial benchmarks object-centric causal spatial reasoning, revealing MLLMs score 54% versus human 84% due to ungrounded textual reasoning, fixed by video-simulation framework COW.

Wenxin (Wendy) Ma, Chenlong Wang, Ruisheng Yuan, Hao Chen and 9 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents

AgentOdyssey generates open-ended long-horizon text games to evaluate test-time continual learning, finding top agents far below human performance despite scaling with model strength.

Zheyuan Zhang, Zehao Wen, Bowei Zhang, Andrew Wang and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 8 on Hugging Face · Code ★ 55

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
88%Must read
?Must readVote to see the score

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

Evaluation metrics for AI radiology reports are highly sensitive to reference reporting style, altering model rankings and revealing poor clinical interpretation decoupling.

Daniel P Jeong, Charles Q Li, Hossein Hosseiny, Nitya M Bhalla and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
86%Must read
?Must readVote to see the score

NORA: Evaluating Grounded Reasonableness in Visual First-person Normative Action Reasoning

NoRA evaluates visual first-person normative reasoning by requiring models to generate actions with fact-reason-action support graphs, revealing current VLMs struggle to bind correct justifications to actions.

Sichao Li, Sai Ma, Zhuang Li, Daniel Kilov and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
80%Must read
?Must readVote to see the score

FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting

FinReasoning hierarchically benchmarks LLMs on financial research via semantic consistency, data alignment, and insight, revealing closed-source models suit core reasoning, open-source models lack consistency, and financial models lack auditing skills.

Yiyun Zhu, Yidong Jiang, Ziwen Xu, Yinsheng Yao and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Human-AI Teaming Through the Lens of Calibration

Calibrated human-AI teaming shows combination methods lose human calibration, while delegation preserves predictor calibration but requires an unattainably precise rejector.

Eric Nalisnick, Chi Zhang, Chengxin Qian, Yixin Wang

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Leveraging Latent Visual Reasoning in Silence

Latent visual reasoning enhances multimodal training despite being largely unused at inference; attention-based reinforcement learning preserves its benefits by promoting latent-text interaction during training.

Dongyao Zhu, Zhen Wang, Xi Xiao, Han Jiang and 6 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
91%Must read
?Must readVote to see the score

Jointly Reinforcing Diversity and Quality in Language Model Generations

DARLING uses a learned partition function to jointly optimize language model response quality and semantic diversity via reinforcement learning, improving both quality and novelty across creative and math benchmarks.

Tianjian Li, Yiming Zhang, Ping Yu, Swarnadeep Saha and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 25 on Hugging Face · Code ★ 61

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
89%Must read
?Must readVote to see the score

WorldMemArena: Evaluating Multimodal Agent Memory Through Action–World Interaction

WorldMemArena evaluates multimodal agent memory through an action-world loop, showing writing and storage improvements do not guarantee performance and harness-based memory remains costly and unreliable.

Chengzhi Liu, Yuzhe YANG, Sophia Xiao Pu, Yepeng Liu and 15 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 13 on Hugging Face · Code ★ 29

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
80%Must read
?Must readVote to see the score

RATS! Patches Talk Through Registers: Emergent Parts in Register Attention Transformers

RATS decomposes vision transformers' classification token into learnable register tokens that spontaneously specialize into object parts, improving segmentation by up to 12 mIoU.

Timing Yang, Predrag Neskovic, Jansen Seheult, Wenchao Han and 3 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
89%Must read
?Must readVote to see the score

Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception

Behavioral geometric supervision aligns video foundation models with human social perception by matching embedding geometry to human similarity judgments, nearly tripling V-JEPA 2.1 performance past language baselines and developing interpretable social attributes with zero-shot transfer.

Kathy Garcia, Leyla Isik

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5
91%Must read
?Must readVote to see the score

STRAND: Sequence-Conditioned Transport for Single-Cell Perturbations

STRAND predicts single-cell transcriptional responses to perturbations by conditioning on regulatory DNA sequence, enabling zero-shot inference across ~95% of the genome with improved discrimination and transfer performance.

Boyang Fu, Sameer Gabbita, George Dasoulas, xiang lin and 4 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
89%Must read
?Must readVote to see the score

Principia: Relational Physics Tests for Video Models

Principia benchmarks video generators via calibration-independent relational physics consistency across eight Newtonian phenomena, finding top models score below 0.42 despite high VBench ratings.

Varun V Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026 · ▲ 18 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 4/5
72%Highly rated
?Highly ratedVote to see the score

SyncLight: Single-Edit Multi-View Relighting

SyncLight propagates single-view lighting edits across arbitrary uncalibrated multi-view captures via a latent bridge-matched diffusion transformer, enabling consistent high-fidelity relighting without camera poses.

David Serrano-Lozano, Anand Bhattad, Luis Herranz, Jean-Francois Lalonde and 1 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
80%Must read
?Must readVote to see the score

Towards Reliable LLM Evaluation: Correcting the Winner’s Curse in Adaptive Benchmarking

SIREN corrects adaptive LLM evaluation's winner's curse via splitwise selection and bootstrap inference for reliable procedure-level performance estimates.

Yang Xu, Jiefu Zhang, Haixiang Sun, Zihan Zhou and 2 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
89%Must read
?Must readVote to see the score

Inertia-1: An Open Exploration of Wearable Motion Foundation Models

Inertia-1 explores wearable motion foundation models via 18.2M hours of accelerometer data, yielding state-of-the-art recipes and open design principles for diverse sensing tasks.

Zongzhe Xu, Aakarsh Anand, Sarah Jiang, Chuntung Zhuang and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 35

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
80%Must read
?Must readVote to see the score

dRAE: Representation Autoencoder with Hyper-Spherical Codes

Hyper-spherical quantization decouples semantics from magnitude via angular routing to prevent codebook collapse, enabling scalable discrete representation autoencoders with full codebook usage and high-fidelity reconstruction.

Tianren Ma, Lin Long, Chuyan Chen, Mu Zhang and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 10 on Hugging Face · Code ★ 10

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
83%Must read
?Must readVote to see the score

AI Evaluation Should Require Standardized Item-Level Data Releases

Standardized item-level benchmark releases should become AI evaluation infrastructure because aggregate scores obscure validity failures; OpenEval archives 10M responses to enable auditability and recover benchmark validity evidence.

Han Jiang, Susu Zhang, Dongyao Zhu, Yuzhuo Bai and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
88%Must read
?Must readVote to see the score

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation

PhyMotion evaluates human video motion via physics-simulated 3D trajectory rewards across kinematics, contact, and dynamics, improving RL post-training realism by +68 Elo.

Yidong Huang, Zun Wang, Han Lin, Dong-Ki Kim and 5 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 6 on Hugging Face · Code ★ 49

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
71%Highly rated
?Highly ratedVote to see the score

Total Variation Rates for Riemannian Flow Matching

Nonasymptotic total variation analysis of Riemannian flow matching bounds sampling error by discretization and learning terms via curvature-aware differential inequalities. Explicit polynomial iteration complexities follow on hyperspheres and SPD manifolds.

Yunrui Guan, Krishnakumar Balasubramanian, Shiqian Ma

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 2/5
medium 4/10
strict 1/5
88%Must read
?Must readVote to see the score

The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory

Proposed E-P-R framework diagnoses AI agents consuming conflicting memory via entry-propagation-recovery, finding a compliance trap where early adoption collapses success.

Yixiong Chen, Xinyi Bai, Alan Yuille

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
80%Must read
?Must readVote to see the score

Personal Visual Memory from Explicit and Implicit Evidence

VisualMem adds structured personal visual memory to text backends, improving personalized agent recall of explicit and implicit visual evidence.

Viet Nguyen, Thao Nguyen, Vishal Patel, Yuheng Li

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Re-evaluating Confidence Remasking in Masked Diffusion Language Models

Post-hoc confidence remasking in masked diffusion language models offers little benefit under standard decoding and worsens diversity collapse under stochastic sampling, showing setting-dependent gains.

Stipe Frković, Metod Jazbec, Dan Zhang, Christian Andersson Naesseth and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 2/5
83%Must read
?Must readVote to see the score

Heavy-Tailed Flow Matching via Random Clocks

HTFM models heavy-tailed sources as mixtures of clock-conditioned Gaussians to improve mode coverage, sample quality, and tail recovery over Gaussian flow matching while enabling tail calibration via clock laws.

Zhouhao Yang, Yezhen Wang, Kenji Kawaguchi, Vladimir Braverman and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

On the Rademacher Complexity of Graph Neural Networks: Unifying Expressivity and Geometry

Rademacher complexity bounds unify GNN expressivity and input geometry through equivalence classes, covering numbers, and Wasserstein robustness.

Martin Carrasco, Caio Deberaldini Netto, Ehimare Okoyomon, Aneeqa Mehrab and 2 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 2/5
medium 5/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

A Multimodal Benchmark for Evaluating Cause-of-Death Inference Using Child Health and Mortality Data

This paper introduces a multimodal benchmark for cause-of-death inference in child mortality data, showing zero-shot language models synthesize unstructured medical evidence differently than supervised baselines.

Junhe Yang, Soumyakanti Pan, Hyun Seung Lim, YUE CHU and 17 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Thinking in Boxes: 3D Editing in Real Images Made Easy

The method treats 3D box pairs as structured transformation specs for precise real-image editing, outperforming state-of-the-art on large 3D edits.

Pradhaan Bhat, Naveen Chandra R, Rishubh Parihar, Vaibhav Vavilala and 3 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Taking the Road Less Scheduled with Adaptive Polyak Steps

Adaptive Polyak step sizes for Schedule-Free SGD and Adam compute iteration-wise learning rates from losses and gradients, achieving anytime convergence without tuning base rates or horizons.

Dimitris Oikonomou, Matthew Buchholz, Yuen-Man Pun, Robert Gower and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Graph Cascades: Contagion-Based Mesoscopic Rewiring for Structure-Aware Graph Machine Learning

Graph Cascades uses contagion diffusion to build auxiliary edges in linear time, boosting GNN and graph transformer accuracy on heterophilic and high-degree graphs while failing on regular low-degree graphs.

Meher Chaitanya Pindiprolu, My Le, Luana Ruiz

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 2/5
86%Must read
?Must readVote to see the score

MoSE3: Learning World-Space SE(3) at Every Pixel

MoSE3 predicts dense per-pixel world-space SE(3) motion from monocular video via point tracks and rigidity embeddings, achieving state-of-the-art 6-DoF estimation and 3D tracking.

Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 3/5
80%Must read
?Must readVote to see the score

Robust Inference-Time Steering of Protein Diffusion Models via Embedding Optimization

EmbedOpt steers protein diffusion by optimizing conditional embeddings rather than atomic coordinates, improving robustness and cryo-EM fitting performance.

Minhuan Li, Jiequn Han, Pilar Cossio, Luhuan Wu

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5