Good Papers

Showing papers from University of Waterloo Show all papers

86%Must read
?Must readVote to see the score

X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization

X-Tree learns reusable hierarchical skills from agent trajectories and improves success rates up to 5.8% across web and science benchmarks.

Sitao Cheng, Xunjian Yin, Zhiyuan Sun, Yuxuan Li and 3 more

Published Sep 26, 2026 · 0 citations · ▲ 76 on Hugging Face · Code ★ 2

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 0/5
78%Highly rated

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Program-as-Weights compiles natural-language specs into compact local neural adapters that match large-model prompting with far less memory and faster offline execution.

Wentao Zhang, Liliana Hotsko, Woojeong Kim, Pengyu Nie and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published Jul 2, 2026 · 0 citations · ▲ 307 on Hugging Face · Code ★ 359

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Tuna-2 replaces vision encoders with patch embeddings for end-to-end pixel-space multimodal understanding and generation, achieving state-of-the-art results that outperform encoder-based designs at scale.

Zhiheng Liu, Weiming Ren, Xiaoke Huang, Shoufa Chen and 11 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 70 on Hugging Face · Code ★ 756

– ReadersNo votes yet. 1 from authors or colleagues not counted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Surprises in Proper Positive-Only Learning

Shai Ben-David, Farnam Mansouri, Anay Mehrotra, Manolis Zampetakis

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
67%Highly rated
?Highly ratedVote to see the score

Settling the Sample Complexity of Deterministic Agnostic PAC Learning

Shai Ben-David, Steve Hanneke, Farnam Mansouri, Amirreza Shaeiri

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 0/5
medium 1/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

FireMPC: A Multi-Source Pan-Canadian Wildfire Benchmark Revealing Spatiotemporal Generalization Gaps

Zhengsen Xu, Sibo Cheng, Lanying Wang, Aryan Sharma and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Learning from Disagreement: Maximum Divergence Knowledge Distillation

Aref Jafari, Parsa Ashrafi Fashi, Mehdi Rezagholizadeh, Hanieh Asadi Golmankhaneh and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

TeMPO: Frame-Causal Token Compression for Efficient Video Large Language Model

Yingxin Lai, Bo Xu, Yun-ze Pan, Zhiliang Zhu and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Rethinking LoRA Initialization for Robust Asymmetric Learning Rates

Disen Liao, Yaoliang Yu

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ASH: Agents that Self-Hone via Embodied Learning

Benjamin Schneider, Xavier Schneider, Victor Zhong, Sun Sun

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A Private Empirical Defense Against Privacy Audits

Saloni Modi, Srivi Balaji, Yusong Zhu, Gautam Kamath and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Hit Expansion via Localized Exploration of Synthesizable Chemical Space

Walter Virany, Yidong Jin, Andrew Lian, Dmytro Shevchuk and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

LaSA-Net: A Language-Guided Network for Outdoor Generalized 3D Referring Expression Segmentation

Lingfei Ma, Bin Liu, Wen Li, Wentao Sun and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Position: Fair Representations Cannot Hold What They Promise

Shai Ben-David, Tosca Lechner, Ruth Urner

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler Chang, Catherine Arnett, Abdelrahman Sadallah, Abdelrahman Eldesokey and 36 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

RTEB: An Overfitting-Resistant Benchmark for Embedding Model Evaluation

Sahil Verma, Minghan Li, Andrew Gaut, Yujie Qian and 14 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Deformable 2D Gaussian Splatting

Huibin Li, Chul Min Yeum

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

Post-hoc Selective Classification for Reliable Synthetic Image Detection

ReSIDe applies post-hoc selective classification to synthetic image detectors by aggregating layer-wise confidence scores via preference optimization, reducing AURC by up to 69.55% under covariate shift.

Kaixiang Zheng, Jacob Seidman

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
83%Must read
?Must readVote to see the score

Stochastic Reconfiguration as Statistical Filtering for Overparameterized Neural Quantum States

Stochastic reconfiguration acts as ridge regression filtering finite-sample noise in overparameterized neural quantum states, and multi-shift averaging lowers validation risk and variance.

Tak Hur

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 2/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

MedMisBench: Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

MedMisBench reveals LLM medical accuracy collapses from 71% to 38% under misleading context, exposing a critical evaluation blind spot around epistemic resilience.

Hongjian Zhou, Xinyu Zou, Jinge Wu, Sean Wu and 18 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5
72%Highly rated
?Highly ratedVote to see the score

Complexity of Classical Acceleration for $\ell_1$-Regularized PageRank

Standard FISTA is asymptotically worse than ISTA for ℓ1-regularized PageRank, though over-regularized objectives with confinement yield accelerated bounds plus boundary overhead.

Kimon Fountoulakis, David Martínez-Rubio

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 1/5
medium 4/10
strict 3/5
89%Must read
?Must readVote to see the score

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation

A unified perturbation framework shows modern pairwise-comparison leaderboards are non-robust, as sub-1% targeted changes alter rankings and confidence intervals.

Hosna Oyarhoseini, Jimmy Lin, Amir-Hossein Karimi

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
92%Must read
?Must readVote to see the score

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

TRL-Bench standardizes cross-paradigm evaluation of tabular encoders via shared representation-level probes, finding encoder quality is task-specific and best pipelines combine capability-matched specialists.

Wei Pang, Xiangru Jian, Hehan Li, Zhixuan Yu and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 54 on Hugging Face · Code ★ 10

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
78%Highly rated
?Highly ratedVote to see the score

Certification from Examples is Hard for Circuits and Transformers under Minimal Overparametrization

Certification becomes exponentially hard for slightly overparameterized depth-2+ circuits and constant-overhead transformers, requiring exponentially large example sets and allowing imperfect models to hide errors.

Artur Back de Luca, Kimon Fountoulakis

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 2/5
89%Must read
?Must readVote to see the score

Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction

Direct corpus interaction uses terminal tools to search raw corpora directly, bypassing fixed retrieval interfaces and substantially outperforming sparse, dense, and reranking baselines on agentic search benchmarks.

Zhuofeng Li, Haoxiang Zhang, Cong Wei, Pan Lu and 14 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 125 on Hugging Face · Code ★ 408

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
86%Must read
?Must readVote to see the score

VistaQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence

VistaQA benchmarks joint visual question answering and pixel-level evidence grounding across 1,157 expert-curated samples, revealing state-of-the-art models achieve limited alignment between answers and visual evidence.

Mozhgan Nasr Azadani, Yimu Wang, Yongpeng Zhu, Lihong Chen and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5