Good Papers

Showing papers from Microsoft Show all papers

86%Must read
?Must readVote to see the score

TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment

TeacherGRPO aligns teachers to student distributions via reinforcement learning to overcome reasoning distillation's Gap Curse and improves student performance.

Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng and 4 more

Published Aug 20, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

ECHO: Terminal Agents Learn World Models for Free

ECHO trains terminal agents to predict environment responses for dense supervision, doubling GRPO pass@1 on TerminalBench-2.0.

Vaishnavi Shrivastava, Ahmed Awadallah, Dimitris Papailiopoulos

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
45%Niche pick
?Niche pickVote to see the score

Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events

Isabella Thiel, Juan M. Bello-Rivas, Yannis Kevrekidis, Themis Sapsis

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Rethinking Language Model Scaling under Transferable Hypersphere Optimization

Liliang Ren, Yang Liu, yelong shen, Weizhu Chen

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Test-Time Learning with an Evolving Library

Weijia Xu, Alessandro Sordoni, Chandan Singh, Zelalem Gero and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

The $1/\mathcal{W}$ Law: Context Length is the Dominant Energy Lever in LLM Inference Fleets

Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CAPO: A Primal-Dual Framework for Constraint-Aware Prompt Optimization

Victor Ye Dong, Reid Pryzant, Yi Liu, Jian Jiao

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AdShot: Benchmarking Multimodal Large Language Models for Video Advertisement Clipping

Wen Xie, Om Rastogi, Sai S Rangoju, Gijs Overgoor and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ALAM: Algebraically Consistent Latent Transitions for Vision-Language-Action Models

Zuojin Tang, Haoyun Liu, Xinyuan Chang, Changjie Wu and 10 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

RoutingBench: Can Agentic Routing Analysis Scale to Production Datacenter Networks?

Wenlong Ding, Zhixiong Niu, Jianan Yang, Fajun Zhang and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

LogT: Logically Think with Images for Visual Search

Yanjun Fu, Quanzeng You, Jiadong Guo, Yujie Lu and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Agent$^2$ RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?

Wanyi Chen, Xiao Yang, Xu Yang, Tianming Sha and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

PreDiff: Sequential Recommendation by Denoising Preference Distributions

Yaoqi Chen, Jianjin Zhang, Qi Chen, Weihao Han and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Synthetic Tasks for Training AutoResearch

Ziyang Cai, Seyyedamirhossein Saeidi, Harkirat Singh Behl

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Synthetic Web: Benchmarking Language Agents under Adversarial Search Ranking

Shrey Shah, Levent Ozgur

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Multi-Turn RL Makes Small Language Model Competitive for Optimization Modeling

Xinzhi Zhang, Zeyi Chen, Humishka Zope, Hugo Barbalho and 5 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Relaxed On-Policy Distillation: Selective Credit Allocation for Scaling Reasoning Efficiently

Jongwoo Ko, Sara Abdali, Young Jin Kim, Tianyi Chen and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

STAR-Math: Multi-Agent Mathematical Reasoning under Persistent Meta-Strategic Supervision

Jiaao Wu, Xian Zhang, Hanzhang Liu, Sophia Zhang and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

ProgramBench: Can Language Models Rebuild Programs From Scratch?

John Yang, Kilian Lieret, Jeffrey Ma, Parth Thakkar and 8 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CSBench: A Comprehensive Benchmark for Evaluating Project-Level System Construction in Computer Science

Hongli Yu, Huan-ang Gao, Botian Wang, Hanlin Wu and 11 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Unified Noise Steering for Efficient Human-Guided VLA Adaptation

Junjie Lu, Xinyao Qin, Yuhua Jiang, Kaixin Wang and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

TreemapMix: Dirichlet-Controlled Multi-Image Augmentation for Probability and Ordinal Supervision

Ejafa Bassam, Konstantin Garbers, Yingsheng Geng, Dalton Jens and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Active Corpus Selection for Training Subgraph Retrievers Using OOD Queries

Pritish Chakraborty, Aditya Singh, Indradyumna Roy, Lokesh N and 7 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

EmpathyChat: Structured Cognitive Reasoning in Empathetic Spoken Dialogue

Dingdong WANG, Shujie LIU, Jinyu Li, Yuxuan Hu and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

fxBench: Evaluating and Understanding Formula Suggestions in Spreadsheets

Sanket Mhatre, Sumit Gulwani, Vu Le, Yasharth Bajpai and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Cascaded Sparse Autoencoders LearnMulti-Level Visual Concepts in Multimodal LLMs

Yusong Zhao, Hengyi Wang, Tanuja Ganu, Akshay Nambi and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
70%Highly rated
?Highly ratedVote to see the score

Bigger Isn’t Better: Why the Indiscriminate Scaling of Foundation Models Can’t Solve Biology

Kathryne Metcalf, Lorin Crawford, Mary L Gray, Kevin K Yang and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 4 of 20 reviewers recommend it
lenient 2/5
medium 1/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

The Era of Agentic Organization: Learning to Organize with Language Models

Zewen Chi, Li Dong, Qingxiu Dong, Yaru Hao and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ContractBench: Can LLM Agents Preserve Observation Contracts?

Jicheng Wang, Yifeng He, Zili Wang, Hanwen Xing and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Exploring Advertising Manipulation in Diffusion Image Generation

Tianshi Che, Yang Zhou, Yushan Mu, Zeru Zhang and 7 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Mining Logic under Uncertainty: Probabilistic Soft Logic with Energy-Based Inference for Chain-of-Thought Verification

Jiang Yu, Jinlong Tian, Kewei Cheng, Yue He and 6 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

SSD: Shell-Guided Spherical Diffusion for Molecular Geometry Generation

Yun-Yen Chuang, Chen-Sheng Gu, Hung-Min Hsu, Kevin Lin and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CoTrek: Toward Scalable On-Policy Distillation for Long Chain-of-Thought Reasoning

Heng Zhang, Chengyu Zhou, Jiajun Wu, Estella Liu and 7 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

PolyMind: Exploring Width Scaling for Reflective Reasoning in Language Agents

Heng Zhang, Chengyu Zhou, Jiajun Wu, Estella Liu and 7 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

HoloCode: A Code-Centric Multi-Agent Framework for Image-to-3D Scene Generation

Hanlin Chen, Chung-Ching Lin, Yuyang Zhao, Dongyue Lu and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

When Does Subspace Direction Matter for LoRA? Regime Analysis of the Magnitude Principle in Few-Shot Adaptation

Nischal Subedi, Cencheng Shen, Peng Zhao

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A Cross-Interaction Neural Architecture for Submodular Functions

SOUTRIK SARANGI, Aditya Singh, Vansh Maheshwari, Abir De

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Data Auctions for Retrieval Augmented Generation

Minbiao Han, Seyed A Esmaeili, Michael Albert, Haifeng Xu

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Firefly: Illuminating Verified Real-world Tool Call Data Generation

Yuxuan Lu, Ziyi Wang, Yingzhou Lu, Yisi Sang and 11 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs

Ming-Hao Hsu, Yuxuan Hu, Shujie LIU, Jinyu Li and 1 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Conditional Evaluation of Language Models with Cheap Auxiliary Signals

LACE estimates conditional LLM accuracy profiles using cheap auxiliary signals via local centering to achieve calibration-free, unbiased, locally optimal semi-supervised evaluation.

Zhi Zhang, Lingfeng Lyu, Yue Kang, Doudou Zhou

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
83%Must read
?Must readVote to see the score

Swift Sampling: Selecting Temporal Surprises via Taylor Series

Swift Sampling uses Taylor-series projections of visual feature trajectories to select temporally surprising frames, cutting overhead by 30x while boosting long-video accuracy up to 12.5 points.

Dahye Kim, Bhuvan Sachdeva, Karan Uppal, Naman Gupta and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing

SafeMoE isolates unsafe knowledge into domain-specific LoRA experts and routes them with a lightweight gating network to improve safe response rates by over 20% relative while maintaining informative outputs.

Maryam Hashemzadeh Barvarz, Jerry Huang, Minseon Kim, Marc-Alexandre Côté and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Reinforcement World Model Learning for LLM-based Agents

RWML learns action-conditioned world models for LLM agents via self-supervised sim-to-real alignment, outperforming direct task-success RL by up to 6.9 points without expert data.

Xiao Yu, Baolin Peng, Ruize Xu, yelong shen and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 28 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

OATS: Online Data Augmentation for Time Series Foundation Models

OATS dynamically generates training-stage-specific synthetic data guided by valuable samples via diffusion, consistently outperforming static augmentation for time series foundation models.

Junwei Deng, Chang Xu, Jiaqi Ma, Ming Jin and 4 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning

DeformGen uses dynamics-based topological augmentation to generate diverse deformable object states and warp trajectories for improved manipulation policy learning.

Zili Lin, Wenyao Zhang, Yuyang Zhang, Zekun Qi and 8 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Video Models Can Reason with Verifiable Rewards

VideoRLVR applies reinforcement learning with verifiable rewards to video diffusion models, improving rule-consistent visual reasoning and cutting training latency 40% via early-step optimization.

Tinghui Zhu, Sheng Zhang, James Yipeng Huang, Selena Song and 4 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 9 on Hugging Face · Code ★ 28

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
89%Must read
?Must readVote to see the score

DPIAgent: Divide, Protocol, Isolate for Agentic Reproduction Test Generation

DPIAgent divides reproduction test generation into isolated diagnosis and test phases with structured handoffs, achieving up to 86.17% success on SWT-Bench Verified.

Hao Liu, Steven Liu, Xin Zhang, Jane Luo and 7 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
89%Must read
?Must readVote to see the score

From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents

P2T uses reference patches as privileged supervision to curate shorter, grounded agent trajectories via bi-objective optimization, improving SWE-bench Pass@1 by up to 10.8 points with ~15% lower inference cost.

Murong Ma, Tianyu Chen, Yun Lin, Shuai Lu and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
89%Must read
?Must readVote to see the score

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents

OpenWebRL enables open online RL for visual web agents, with a 4B model reaching 67% Online-Mind2Web and 64% DeepShop success using minimal initialization data.

Rui Yang, Qianhui Wu, Yuxi Chen, Hao Bai and 6 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 19 on Hugging Face · Code ★ 52

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use

CARL assigns segment-level reinforcement learning credit at tool-use boundaries to teach models when external tools are needed, improving accuracy by up to 9.7 points while cutting unnecessary calls by 53%.

Abhijit Kumar, Zoey WU, Mohit Suley

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

The FACTS Leaderboard benchmarks large language model factuality across multimodal, parametric, search, and grounding tasks via automated judges.

Aileen Cheng, Alon Jacovi, Amir Globerson, Ben Golan and 36 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 1/5
80%Must read
?Must readVote to see the score

Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds

Intrinsic Muon extends norm-constrained matrix optimization to Riemannian manifolds via canonical intrinsic norms, yielding closed-form updates and convergence rates independent of factor conditioning.

Yibang Li, Bihari L Pandey, Ravi Sah, Andi Han and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

TwinRouterBench introduces static and live dynamic tracks to benchmark LLM routing at agent step-level using deterministic scoring and live execution on SWE-bench.

Pei Yang, Wanyi Chen, Tongyun Yang, Pengbin Feng and 11 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding

XL-DocBench introduces a human-verified benchmark for extra-long professional document understanding spanning thousands of pages with multi-page evidence, showing current systems still struggle with long-context structured reasoning.

Hongchen Wei, Yuanzhe Wang, Bei Liu, Yifan Yang and 8 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
86%Must read
?Must readVote to see the score

DocAtlas: Long-Document Understanding as Mutable-State Interaction

DocAtlas treats long-document understanding as a mutable-state interaction process via a document harness with search, memory, and review tools, reaching 71.4% on MMLongBench-Doc and boosting a 4B VLM to 63.7% via reinforcement learning.

Hongchen Wei, Yuanzhe Wang, Bei Liu, Yifan Yang and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 12

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Rethinking Vector Field Learning for Generative Segmentation

Flow matching for generative segmentation suffers gradient vanishing and trajectory traversing, causing slow convergence and poor class separation; reshaping the velocity field with distance-aware corrections and Kronecker-based category encoding narrows the gap with discriminative specialists.

Chaoyang Wang, Yaobo Liang, Boci Peng, Fan Duan and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

92%Must read
?Must readVote to see the score

MorphoHELM: A Comprehensive Benchmark for Evaluating Representations for Microscopy-Based Morphology Assays

MorphoHELM benchmarks microscopy representation methods across batch effects, finding classic computer vision strategies outperform deep learning across settings and revealing trade-offs between models.

Emre Hayir, Lorin Crawford, Alex X Lu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
76%Highly rated
?Highly ratedVote to see the score

Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory

Memory Grafting uses frozen hidden states from a grafting model as conditional n-gram memory for language models, improving average benchmarks to 53.86 versus 52.43 for vanilla Engram at 2.8B scale with minimal overhead.

Runxi Cheng, Yuchen Guan, Yongxian Wei, Qianpu Sun and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
80%Must read
?Must readVote to see the score

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning

G2PO transforms agent trajectories into state-transition graphs to reduce variance and improve credit assignment, outperforming GRPO by up to 22.2% on long-horizon benchmarks.

Yunan Wang, Minghui Song, Zihan Zhang, Shaohan Huang and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
86%Must read
?Must readVote to see the score

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

STARE analyzes token-level entropy dynamics under GRPO, identifies a credit assignment mismatch, and uses surprisal-guided advantage reweighting to stabilize policy entropy, improving AIME accuracy by 4-8%.

HAIPENG LUO, Qingfeng Sun, Song-Li Wu, Can Xu and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 12 on Hugging Face · Code ★ 24

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text

LensVLM lets VLMs scan compressed rendered text and selectively expand only relevant regions via learned tools, maintaining near-full accuracy at 4.3x compression and outperforming baselines up to 10.1x across text QA benchmarks.

Roy Xie, Dan Friedman, Donghan Yu, Bowen Pan and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

Medmarks introduces 30 open-source medical benchmarks evaluating 61 LLMs, finding frontier reasoning models lead, proprietary models are more token-efficient, medical fine-tuning helps, and smaller models show answer-order bias.

Benjamin Warner, Ratna S Grandhi, Max Kieffer, Aymane Ouraq and 31 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
83%Must read
?Must readVote to see the score

Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task

LLMs diverge from human goal selection in self-directed learning by exploiting single solutions with low variability instead of gradual diverse exploration.

Gaia Molinaro, Dave August, Danielle Perszyk, Anne Collins

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

AI Evaluation Should Require Standardized Item-Level Data Releases

Standardized item-level benchmark releases should become AI evaluation infrastructure because aggregate scores obscure validity failures; OpenEval archives 10M responses to enable auditability and recover benchmark validity evidence.

Han Jiang, Susu Zhang, Dongyao Zhu, Yuzhuo Bai and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
70%Highly rated
?Highly ratedVote to see the score

Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference

This paper models parallel inference-time reasoning via particle filtering, deriving non-asymptotic guarantees and fundamental limits for sequential Monte Carlo with process reward models.

Noah Golowich, Fan Chen, Dhruv Rohatgi, Raghav Singhal and 3 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning

SWAP assigns step-level length penalties based on reasoning contribution, cutting chain-of-thought length by 64.3% and boosting accuracy 5.7% over base models.

Xintong Li, Sha Li, Rongmei Lin, Hongye Jin and 9 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction

ViDiHand leverages pretrained video diffusion models to reconstruct 4D hand poses directly from full egocentric video without detectors, substantially outperforming prior methods on ARCTIC, HOT3D, and HOI4D.

Yuxi Wang, Chengkai Jin, Yufei Liu, Wenqi Ouyang and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 9 on Hugging Face · Code ★ 187

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

ScreenSearch: Uncertainty-Aware OS Exploration

ScreenSearch combines structural screen retrieval with ambiguity-aware PUCT search to explore desktop OS states, collecting over 1M screenshots across 11 apps and showing ambiguity reduction alone is insufficient for exploration.

Michael Solodko, Justin Wagle

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
86%Must read
?Must readVote to see the score

Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors

A clustering-based divergence method measures gaps between real and simulated user behaviors, finding large, family-dependent discrepancies reducible by combining complementary simulators.

Shuhaib Mehri, Philippe Laban, Sumuk Shashidhar, Marwa Abdulhai and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph

This paper formalizes test-time scaling as an optimizable multi-LLM collaboration graph and proposes Agent-REINFORCE to search compute-optimal architectures under budget constraints. Experiments show it outperforms baselines in efficiency and finds graphs balancing accuracy with latency.

Fali Wang, Jihai Chen, Shuhua Yang, Runxue Bao and 6 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

RepoLaunch: Automating Build and Management of Code Repositories across Languages and Platforms

RepoLaunch automates cross-language repository builds and testing, achieving 78% build success and enabling fully automated SWE dataset pipelines.

Kenan Li, Rongzhi Li, Linghao Zhang, Qirui Jin and 16 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Securing AI Agents with Information-Flow Control

Information-flow control secures AI agents via Fides, a planner with deterministic confidentiality and integrity tracking that completes diverse AgentDojo tasks with guarantees.

Manuel Costa, Boris Köpf, Aashish Kolluri, Andrew Paverd and 5 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026 · ▲ 1 on Hugging Face · Code ★ 118

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic

ETL monitors perception-based autonomous systems directly in learned embedding spaces via distance-based temporal logic predicates, enabling reliable specification of high-level visual behaviors with conformal calibration.

Parv Kapoor, Abigail Hammer, Ashish Kapoor, Karen Leung and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
88%Must read
?Must readVote to see the score

Get a GRIP, this will be a long TRIP: A Quantifiable Long-Range Framework for Verifying Over-squashing

Introducing verifiable axioms for long-range graph benchmarks, this work proposes TRIP/GRIP to construct provably long-range tasks with closed-form per-range error bounds and audits existing benchmarks.

Ferran Hernandez Caralt, Simon Heilig, Adrián Bazaga, Asja Fischer and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Tracing Agentic Failure from the Flow of Success

OAT trains neural controlled differential equations on successful agent trajectories to detect failure steps without failure annotations, outperforming prompting baselines by up to 20% F1 with 200-5000x speedup.

Samuel (Min-Hsuan) Yeh, Yiwen Zhu, Shaleen Deep, Sharon Li

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 16 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
91%Must read
?Must readVote to see the score

GeneZip: Region-Aware Compression for Long Context DNA Modeling

GeneZip uses region-aware compression to achieve high base-pairs-per-token ratios, improves DNA modeling benchmarks, and enables 128K-context training on limited hardware.

Jianan Zhao, Xixian Liu, Zhihao Zhan, XINYU YUAN and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
80%Must read
?Must readVote to see the score

M$^\star$: Every Task Deserves Its Own Memory Harness

M* evolves task-specific memory programs via reflective code search to outperform fixed-memory agents across diverse benchmarks. Evolved harnesses develop structurally distinct mechanisms per domain, showing specialization beats general-purpose memory.

Wenbo Pan, Shujie LIU, Xiangyang Zhou, Xianlong Wang and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Vermeer: Autoregressive generative modeling of microscopy predicts protein localization

Vermeer is an autoregressive generative model that predicts protein localization microscopy from sequences and cell landmarks, enabling zero-shot transfer across imaging conditions.

Sandeep Kambhampati, Eric Zimmermann, Emre Hayir, Kevin K Yang and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

NGDB-Zoo: Towards Efficient and Scalable Neural Graph Databases Training

NGDB-Zoo improves neural graph database training via operator-level scheduling and semantic augmentation, achieving 1.8, 6.8x throughput without I/O stalls.

zhongwei xie, Jiaxin Bai, Shujie LIU, Haoyu Huang and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 2/5
medium 6/10
strict 1/5
Show 20 more papers