Good Papers

Showing papers from University of Toronto Show all papers

89%Must read
?Must readVote to see the score

Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers

LLM scorers with equal ranking quality make unstable threshold and preference decisions under candidate reordering, and order-consistency fine-tuning fixes it without harming quality.

Markus Frohmann, Mahdiyar Alavi, Elizabeth Lingg, Navid Rekabsaz

Published Aug 27, 2026 · 0 citations · ▲ 17 on Hugging Face · Code ★ 2

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

Attention Is All You Need

The Transformer replaces recurrence and convolutions with attention, achieving superior translation quality and faster training.

Ashish Vaswani, Noam Shazeer, Niki Jitendra Parmar, Jakob Uszkoreit and 4 more

Published Aug 23, 2025 · 26,828 citations

– ReadersNo votes yet
18/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 21 reviewers recommend it
lenient 5/5
medium 9/11
strict 4/5
45%Niche pick
?Niche pickVote to see the score

Laplacian Heads Improve Transformers by Smoothing Token Representations

Yuchong Zhang, Vardan Papyan

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Sequential Probability Assignment against Smoothed Adversaries with Unknown Base Measure

Ziyi Liu, Dan Roy

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Integrated Imputation-Classification for Supervised Learning with Missing Data

Yue Liu, Ben Liang, Ali Tizghadam, Ilijc Albanese

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

When to Trust a PFN: Detecting Harmful Shift in Tabular Foundation Models

Viet Nguyen, Herman Bergström, Stephan Rabanser, Rahul Krishnan

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Colour me shocked: Exact Molecular Hessians from MLIPs in O(N) time using sparse differentiation!

Luca Thiede, Andreas Burger, Alan Aspuru-Guzik

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DynaSub: Adaptive Subgrouping for Scalable Representation Learning

Tina Behrouzi, Sana Tonekaboni, Rahul Krishnan, Anna Goldenberg

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SuperSycophantic: Stress-Testing Frontier LLMs from Single- to Multi-Turn Sycophancy

Terry J Zhang, Oscar S Yasunaga, Wenyuan Jiang, Jessica Bo and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

FedLoVA: Value-Only Aggregation for Federated LoRA Fine-Tuning of Large Language Models

Ensieh Khazaei, Baturalp Buyukates, Dimitrios Hatzinakos

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

On the Selectivity of Generative Models in Structure-Based Drug Design

Ella Miray Rajaonson, Jungyoon Lee, William Chau, Alan Aspuru-Guzik and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

STAR-Math: Multi-Agent Mathematical Reasoning under Persistent Meta-Strategic Supervision

Jiaao Wu, Xian Zhang, Hanzhang Liu, Sophia Zhang and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Plan2Sense: Open-World Task Planning in Epistemic States via Interleaved Ontic and Sensing Actions

Xiaotian Liu, Armin Toroghi, Jiazhou Liang, Ali Pesaranghader and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Runtime Verification of Multiple Natural Language Criteria for Agent Governance

Silviu Pitis, Parand A. Alamdari, Jessica Tang, Toryn Klassen and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Towards Self-Supervised, Generalizable and Decomposable 4D Driving Scene Reconstruction

James Tu, Anqi Joyce Yang, Jingkang Wang, Sivabalan Manivasagam and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

What should post-training optimize? A test-time scaling law perspective

Muheng Li, Jian Qian, Wenlong Mou

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Hit Expansion via Localized Exploration of Synthesizable Chemical Space

Walter Virany, Yidong Jin, Andrew Lian, Dmytro Shevchuk and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Think about how AIs think about themselves

Raymond Douglas, Jan Kulveit, Ondřej Havlíček, Theia Pearson-Vogel and 1 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Median-of-Means under Structured Heavy-Tailed Noise: High-Probability Bounds for Clipped Stochastic Optimization

Ahmed El Bajdali, Ohad Shamir, Samuel Horváth, Eduard Gorbunov

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

The Heel of RLVR: Benchmark Glory Should Not Outpace Honest Measurement

Shuo Yang, Chiyu Ma, Kexin Huang, Jinda Lu and 10 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

VerifyThisBench: Joint Evaluation of Code, Specifications, and Proof

Xun Deng, Barış Bayazıt, Si Cheng Zhong, Andreas Veneris and 2 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?

Miaosen Chai, Wang Bill Zhu, Shangshang Wang, Yejia Liu and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Ground False: Uncovering Errors in Formal Mathematics Benchmarks

Marcus Min, One An, Xujie Si, Osbert Bastani

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Literati: Towards Anytime Optimal Shape Generalized Trees via AO*

Nakul Upadhya, Eldan Cohen

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Reinforcement Learning Agents Are Swimmers

Juan Rojas, Jacob Adamczyk, Abhishek Naik, Volodymyr Makarenko and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Foundation Model Informed Acquisition Functions for Molecular Discovery

Qi CHEN, Fabio Ramos, Alan Aspuru-Guzik, Florian Shkurti

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

UTOPI: Efficient Egocentric Long-Video Understanding in AR via User-Guided Token Pre-Compression

ziqi wang, Su Chen, Qiance Tang, Jieyu Lin and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Test-time Scaling of Diffusions with Flow Maps

Flow Map Trajectory Tilting uses flow maps to enable principled diffusion test-time scaling with reward gradients, improving reward ascent and enabling complex image editing via vision-language models.

Amirmojtaba Sabour, Michael Albergo, Carles Domingo i Enrich, Nicholas Boffi and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 0/5
86%Must read
?Must readVote to see the score

DD-Ranking: Rethinking the Evaluation of Dataset Distillation

DD-Ranking reveals dataset distillation gains come from extra evaluation techniques rather than image quality, proposing fair metrics to assess true synthetic dataset value.

Zekai Li, Xinhao Zhong, Samir Khaki, Zhiyuan Liang and 36 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Multi-Marginal Couplings for Metropolis--Hastings

Multi-marginal coupling of Metropolis-Hastings chains via shared-randomness Poisson Monte Carlo improves coalescence rates and reduces meeting times up to 50%.

Truong Buu Phan, Gergely Flamich, Ashish Khisti, Shahab Asoodeh

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention

TurboVGGT enables fast multi-view 3D reconstruction via adaptive alternating attention that balances sparse global and local frame attention while maintaining competitive quality.

David Huang, Guile Wu, Chengjie Huang, Bingbing Liu and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
89%Must read
?Must readVote to see the score

MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models

MedKIT evaluates medical LLM knowledge integration via clinical updates, revealing strong recall but limited relational, compositional, and operational generalization across 12 strategies.

Lukas Thede, Yash Kumar, David Chen, Danielle Bitterman and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5
80%Must read
?Must readVote to see the score

DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning

DriveDreamer-Policy unifies depth generation, video prediction, and motion planning via geometry-aware world representations, achieving 89.2 PDMS on Navsim v1 and 88.7 EPDMS on v2.

Yang Zhou, Xiaofeng Wang, Hao Shao, Letian Wang and 7 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face · Code ★ 60

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

AVIS: Adaptive Test-Time Scaling for Vision–Language Models

AVIS introduces a per-query adaptive policy that jointly scales visual token pruning and reasoning rollouts to improve vision-language model accuracy-compute trade-offs.

Ahmadreza Jeddi, Minh Le, Amirhossein Kazerouni, Hakki Karaimer and 7 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
80%Must read
?Must readVote to see the score

Addressable Memory for Video World Models

Interactive video world models lose visual persistence beyond training horizons because temporal RoPE offsets become out-of-distribution; WorldTrace assigns compressed memory slots virtual in-distribution positions to restore addressability, boosting temporal consistency by 15.5% and episodic recall

Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 16 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
91%Must read
?Must readVote to see the score

P$^{3}$: Joint Program-and-Proof Planning\\ for Verified Code Generation

P³ plans programs and proofs jointly from specifications before elaboration, outperforming sequential baselines by up to 11.2 points on verified generation benchmarks while reducing cost and time.

Zenan Li, Ziran Yang, Peiyang Song, Zhaoyu Li and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
71%Highly rated
?Highly ratedVote to see the score

SGD in Multiclass Logistic Regression: Sequential Learning and Scaling Laws

Multiclass logistic regression learns Gaussian mixtures sequentially by class frequency, producing power-law risk phases and compute-optimal scaling laws.

Konstantinos Tsiolis, Denny Wu, Christos Thrampoulidis, Murat Erdogdu

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Efficient Prediction of Pass@k Scaling in Large Language Models

Standard pass@k scaling laws suffer statistical shortcomings, so a beta-binomial framework and dynamic sampling strategy more accurately predict rare LLM capabilities and risks from limited data.

Joshua Kazdan, Colin Sullivan, Rylan Schaeffer, Youssef Allouah and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
91%Must read
?Must readVote to see the score

Negation Neglect: When models fail to learn negations in training

Fine-tuning LLMs on documents that flag claims as false makes them believe those claims, with belief rates jumping from 2.5% to 88.6%, though local negation phrasing largely prevents it.

Harry Mayne, Lev McKinney, Jan Dubiński, Adam Karvonen and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 4/5
78%Highly rated
?Highly ratedVote to see the score

Vision Transformers Learn Gestalt-Like Figure-Ground Cues from Natural Images

Vision transformers learn Gestalt-like figure-ground cues, surroundedness, convexity, and symmetry for uniform regions, from natural images, with linear probes generalizing zero-shot to artificial stimuli.

Matthias Tangemann, Benjamin Lo, Zygmunt Pizlo, Kaleem Siddiqi and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 2/5
Show 20 more papers