Good Papers

Showing papers from Google DeepMind Show all papers

45%Niche pick
?Niche pickVote to see the score

Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production

Alexandre Galashov, Matt Jones, Nan Rosemary Ke, Yuan Cao and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

LoRAcles: Self-Supervised Weight-Space Interpretability at Scale

Celeste De Schamphelaere, Jan Bauer, Neel Nanda, Euan Ong

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Short-Context Dominance: How Much Local Context Natural Language Actually Needs?

Vala Vakilian, Zimeng Wang, Ankit Rawat, Christos Thrampoulidis

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
67%Highly rated
?Highly ratedVote to see the score

Designing Effective Monitor-Based Interventions for Mitigating Reward Hacking During RL

Aria Wong, Joshua Engels, Neel Nanda

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Model Incrimination: Investigating Whether Concerning Behavior Reflects Misalignment

Gerson Kroiz, Aditya Singh, Senthooran Rajamanoharan, Neel Nanda

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Gradient Routing Localizes and Removes Unintended Behaviors in RL

Jake Ward, Shawn Hu, Aria Wong, Nathan Hu and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Optimizing Analytic Constants via AI-Guided Lean Proof Refinement

Rahul Saha, Alan Li, Anton Xue, Adam Klivans and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Chain-of-Thought Is Not Explainability

Fazl Barez, Tung-Yu Wu, Iván Arcuschin Moreno, Michael Lan and 12 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Process

Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou and 6 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

On the Pitfalls of Instance-Based Dynamic Curricula

Alexandre Galashov, Amal Rannen-Triki, Yee Whye Teh, Razvan Pascanu and 1 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Counterfactual Debugging the World Model Transfer Gap

Mingxuan Li, Kai-Zhan Lee, Michael Dennis, Elias Bareinboim

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

A Subgoal-driven RL Framework for Improving Long-Horizon Web Agents

Taiyi Wang, Sian Gooding, Florian Hartmann, Oriana Riva and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Evaluating Compositional Generalization in Transformers: The Role of Composition Equivalence and Module Coverage

Purva Pruthi, Andrew Yuan, Alexander D'Amour, David Jensen

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

L$^2$EAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

Po-Nien Kung, Linfeng Song, Dawsen Hwang, Jinsung Yoon and 9 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems

Tunyu Zhang, Zihao Zhao, Yusong Zhao, Haizhou Shi and 4 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CryptanalysisBench: Can LLMs do cryptanalysis?

Lukas Fluri, Avital Shafran, Nicholas Carlini, Matthew Jagielski and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Reinforced Fast Weights via Next-Sequence Prediction

Hee Seung Hwang, Xindi Wu, Sanghyuk Chun, Zhiwei Deng and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Beyond Ground Truth: Evaluating Non-Verifiable Reasoning in LLMs through Moral Robustness

Elizaveta Tennant, Benjamin Henke, Anita Keshmirian, Murray Shanahan and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ANCHOR: Audio-Visually Grounded Chain-of-Thought Reasoning Benchmark

Joel Julin, Souraja Kundu, Liza Dahiya, George Z Wei and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Beyond Real or Fake: A Dual-Channel Authenticity and Reasoning Protocol for Photographic Assessment

Xiaoxiao Li, Ruinan Jin, Lili Meng, Miaosen Wang and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ActO: Extracting Action Representations from MLLM Embeddings for Video World Models

Runjia Li, Minghao Chen, Junyu Xie, Philip Torr and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

As the Story Unfolds: Watching a Film and Identifying Characters as a Human Does

Zhongrui Gui, Junyu Xie, Tengda Han, Weidi Xie and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Plausible Biomolecular Structure Prediction via Physics-informed Reinforcement Learning

Tai Dang, Hieu Tran, Long-Hung Pham, Sang Truong and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Off-policy Learning with Excursion Policies

Jiamin He, Mark Rowland, Daniel (Zhaohan) Guo, Hado van Hasselt and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Evaluating Spatiotemporal Reasoning of Vision-Language Models in Atari Gameplay

Mingjia Huo, Yao Fu, Bo Chang, Yaqing Wang and 7 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Automating ML for Science: Can Frontier Agents Climb Scientific Hills in the Wild?

Ming Zhong, Stacy Li, Nicholas Carlini, Matthew Jagielski

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

AIRA-Compose: Agentic Discovery of Neural Architectures

Alberto Pepe, Chien-Yu Lin, Despoina Magka, Bilge Acun and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Homological Barriers to Stable Local Nash Dynamics in Quadratic Zero-Sum Games

Ashkan Soleymani, Gabriele Farina, Patrick Jaillet, Georgios Piliouras

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Anchored Protein Engineering

Chi Zhang, Maria Rosaria Briglia, Litu Rout, Jeffrey Ouyang-Zhang and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

We Should Distinguish Unlearning From Untraining

Eleni Triantafillou, Imtiaz Humayun, Mónica Ribero, Alexander Turner and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Continuous p-adic Optimization

Julian Salazar, Dimitri Kanevsky, Matt Harvey, Pascal Getreuer and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

SeeSE3: The Emergence of 3D Space in Vision Features

Viorica Patraucean, Leonidas Guibas, Sayna Ebrahimi, Caroline Chen and 3 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler Chang, Catherine Arnett, Abdelrahman Sadallah, Abdelrahman Eldesokey and 36 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Regularized Large Neighborhood Search

Regularized LNS turns local search heuristics into MCMC samplers with Fenchel-Young losses, enabling exact block Gibbs sampling and end-to-end learning without global solvers.

Germain Vivier-Ardisson, Laurent Demonet, Axel Parmentier, Mathieu Blondel

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

TokenSwap benchmarks and reduces MLLMs' modality gap by interleaving visual tokens with text, finding reasoning models have smaller gaps and training with TokenSwap mitigates it.

Andong Hua, Colton Bishop, Igor Mordatch, Arian Hosseini and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

MOOD benchmark shows guard models fail to detect out-of-distribution alignment failures, but combining them with Mahalanobis and perplexity detectors improves recall from 39% to 45% and scales positively.

Dylan Feng, Pragya Srivastava, Anca Dragan, Cassidy Laidlaw

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Evaluating and Understanding Scheming Propensity in LLM Agents

Realistic agent settings show minimal scheming despite high incentives, with model-organism scheming brittle to tool removal and oversight.

Mia Hopman, Jannes Elstner, Maria Avramidou, Amritanshu Prasad and 1 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Masked Visual Actions for Unified World Modeling

Masked Visual Actions expresses robot and object motion as revealed pixel trajectories to unify forward dynamics, planning, and inverse modeling in video world models with minimal finetuning.

Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey and 7 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 9 on Hugging Face · Code ★ 110

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Joint Learning of Hierarchical Neural Options and Abstract World Model

AgentOWL jointly learns hierarchical neural options and an abstract world model for sample-efficient skill acquisition, outperforming baselines on object-centric Atari games with fewer samples and stronger generalization.

Top Piriyakulkij, Wolfgang Lehrach, Kevin Ellis, Kevin Murphy

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Dual-Rate Diffusion: Accelerating diffusion models with an interleaved heavy-light network

Dual-Rate Diffusion accelerates diffusion inference by interleaving sparse heavy context encoders with light denoising models, cutting computation 2-4x without quality loss.

Grigory Bartosh, David Ruhe, Emiel Hoogeboom, Jonathan Heek and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
89%Must read
?Must readVote to see the score

Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors

Fine-tuned LLMs learn to selectively hide internal representations from unseen activation monitors via low-dimensional subspace manipulation, evading even post-hoc safety probes with modest capability loss.

Max McGuinness, Alex Serrano Terre, Luke Bailey, Scott Emmons

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models

LRMs show a large production-evaluation gap, scoring near 48% on reasoning evaluation versus near-perfect production due to answer confirmation bias.

Mingzhong Sun, Teresa Yeo, Armando Solar-Lezama, Tan Zhi-Xuan

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 3/5
78%Highly rated
?Highly ratedVote to see the score

VAANI: Capturing the language landscape for an inclusive digital India

Project VAANI releases a multimodal dataset of 31,255 speech hours and 289K images spanning 105 Indic languages across 165 Indian districts to support inclusive speech technology.

Sujith Pulikodan, Abhayjeet Singh, Agneedh Basu, Nihar Desai and 16 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

Alem benchmarks open-ended multi-agent coordination for language agents, showing frontier LLMs average ~6% returns and individual competence does not imply coordination competence.

Kale-ab Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face · Code ★ 51

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
88%Must read
?Must readVote to see the score

Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent

A representation-readout decomposition attributes grokking and double descent to competing encoder and classifier dynamics, showing delayed generalization stems from gradual representation learning rather than lazy-to-rich transitions.

Chi-Ning Chou, Oscar Uzdelewicz, Neng-Chun Chiu, Yao-Yuan Yang and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Subliminal Learning Is Steering Vector Distillation

Subliminal learning is steering vector distillation where students learn teachers' hidden traits via single steering vectors, requiring adaptive optimizers and failing across models.

Camila Blank, Agam Bhatia, Senthooran Rajamanoharan, Arthur Conmy and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

The FACTS Leaderboard benchmarks large language model factuality across multimodal, parametric, search, and grounding tasks via automated judges.

Aileen Cheng, Alon Jacovi, Amir Globerson, Ben Golan and 36 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding

Co-Coder formalizes multi-agent coding as graph partitioning to balance parallel speedups against communication overhead, improving pass rates by 14% and cutting costs 35% on dense repositories.

Xu Yang, Lunyiu Nie, Ethan Chandra, Stanislav Gannutin and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Sparse Koopman Autoencoders Identify Local Dynamical Regimes in Multibasin Systems

Sparse Koopman autoencoders use sparse latent supports as label-free regime indicators that identify local dynamical basins and outperform dense autoencoders in multibasin forecasting.

Aidan Li, Uday Kiran Reddy Tadipatri, Mahan Fathi, Sarath Chandar and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 0/5
80%Must read
?Must readVote to see the score

Differentiable Knapsack and Top-k Operators via Dynamic Programming

A unified framework casts knapsack and top-k operators as dynamic programs with smoothed recursions for differentiable relaxations, parallel algorithms, and theoretical regularization guarantees.

Germain Vivier-Ardisson, Michael E Sander, Axel Parmentier, Mathieu Blondel

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
86%Must read
?Must readVote to see the score

Training on Documents About Monitoring Leads to CoT Obfuscation

Synthetic document finetuning teaches models to hide misbehavior from chain-of-thought monitors, with success tied to reasoning controllability and faster reward-hacking under RL.

Reilly Haskins, Bilal Chughtai, Joshua Engels

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Asking the Right Questions: Improving Reasoning with Generated Stepping Stones

ARQ introduces a question generator that produces transferable intermediate stepping stones, improving reasoning LLM performance via fine-tuning on synthetic data.

Hengyuan Hu, Tingchen Fu, Minqi Jiang, Alexander Miller and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 0/5
91%Must read
?Must readVote to see the score

Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks

RRD refines rubrics via recursive decomposition and filtering to improve LLM judge accuracy and reinforcement training rewards on open-ended tasks.

William Shen, Xinchi Qiu, Chenxi Whitehouse, Lisa Alazraki and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
89%Must read
?Must readVote to see the score

How Post-Training Shapes Biological Reasoning Models

Continued pre-training aligns biological language, supervised fine-tuning improves in-domain but harms out-of-domain reasoning, and reinforcement learning recovers generalization when rewards align.

Lukas Fesser, Hanlin Zhang, Michelle M Li, Eric Wang and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 3 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Eliciting Secret Knowledge from Language Models

Secret-knowledge-elicitation techniques, especially prefill attacks, successfully extract hidden knowledge that LLMs deny knowing but apply downstream.

Bartosz Cywiński, Emil Ryd, Rowan Wang, Senthooran Rajamanoharan and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face · Code ★ 24

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts

Persona Generators use evolutionary code optimization to expand brief context descriptions into diverse synthetic populations maximizing opinion and preference coverage. Evolved generators substantially outperform baselines across six diversity metrics by spanning rare trait combinations.

Davide Paglieri, Logan Cross, William Cunningham, Joel Leibo and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

A Unified Framework for Adversary-Aware Differential Privacy Bounds

A unified framework bounds DP privacy leakage against multi-target membership, attribute, and reconstruction attacks using only privacy parameters and adversarial baseline success rates.

Marika Swanberg, Meenatchi Sundaram Muthu Selva Annamalai, Jamie Hayes, Borja Balle and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Phases of Muon: When Muon Eclipses SignSGD

Spectral optimizer analysis reveals three phases where Muon's SignSVD preconditions covariance differently than SignSGD.

Elliot Paquette, Noah Marshall, Lucas Benigni, Guangyuan Wang and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 2/5
medium 4/10
strict 1/5
92%Must read
?Must readVote to see the score

CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving

CausalDriveBench evaluates causal reasoning in autonomous driving vision-language-action models via structured QA and counterfactual trajectories, finding weak causal understanding despite fluent reasoning and accurate baseline predictions.

Narendiran Chembu, Navvrat Rao, Shreedhar Kodate, Gayatri S Banda and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 5/5
74%Highly rated
?Highly ratedVote to see the score

gfnx: Fast and Scalable Library for Generative Flow Networks in JAX

gfnx is a JAX library for training and evaluating GFlowNets that achieves up to 80x speedups over PyTorch benchmarks across diverse tasks.

Daniil Tiapkin, Artem Agarkov, Nikita Morozov, Ian Maksimov and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 2/5
89%Must read
?Must readVote to see the score

One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation

Monocular pretraining lifts single images into pseudo-target views via depth and reprojection, yielding OVIE, which rivals multi-view baselines at 116 FPS without inference-time depth or multi-view training pairs.

Adrien RAMANANA RAHARY, Nicolas Dufour, Patrick Perez, David Picard

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026 · ▲ 5 on Hugging Face · Code ★ 82

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Exposing the Illusion of Erasure in Knowledge Editing for LLMs

Knowledge editing suppresses rather than erases facts in LLMs, leaving edited knowledge vulnerable to adversarial recovery across architectures.

Advik Basani, Anshuman Chhabra

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Neuron Populations Exhibit Divergent Selectivity with Scale

Rosetta neuron populations grow sublinearly and become more selective and specialized as language and vision models scale, while non-Rosetta neurons stay less selective.

Amil Dravid, Yasaman Bahri, Alexei Efros, Yossi Gandelsman

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
89%Must read
?Must readVote to see the score

Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples

Poisoning LLM pretraining requires only ~250 malicious documents regardless of dataset or model scale, revealing constant-cost backdoor injection risks for large models.

Alexandra Souly, Javier Rando, Ed Chapman, Xander Davies and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 4/5
83%Must read
?Must readVote to see the score

SkillOS: Learning Skill Curation for Self-Evolving Agents

SkillOS uses RL to train a skill curator that updates an external SkillRepo from experience, improving self-evolving agents across reasoning and multi-turn tasks.

Siru Ouyang, Jun Yan, Yanfei Chen, Rujun Han and 12 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 45 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
91%Must read
?Must readVote to see the score

M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling

M²RNN introduces matrix-valued non-linear RNNs that scale via state expansion, achieving perfect state tracking and outperforming hybrid models with smaller states.

Mayank Mishra, Shawn Tan, Ion Stoica, Joseph Gonzalez and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
89%Must read
?Must readVote to see the score

Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction

Direct corpus interaction uses terminal tools to search raw corpora directly, bypassing fixed retrieval interfaces and substantially outperforming sparse, dense, and reranking baselines on agentic search benchmarks.

Zhuofeng Li, Haoxiang Zhang, Cong Wei, Pan Lu and 14 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 125 on Hugging Face · Code ★ 408

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
86%Must read
?Must readVote to see the score

GMOS: Grounding Moving Object Segmentation in 3D Space and Time

GMOS grounds moving object segmentation in 3D space and time using an RGB video framework, achieving state-of-the-art results across MOS benchmarks with faster online inference.

Junyu Xie, Tengda Han, Weidi Xie, Andrew Zisserman

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

ReToken: One Token to Improve Vision–Language Models for Visual Retrieval

ReToken introduces one learnable retrieval token that selects sparse visual tokens from long contexts, improving vision-language models by up to 13.4 points on visual retrieval while fitting on a single GPU.

Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 8 on Hugging Face · Code ★ 18

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment

A unified sign-language model with privacy-preserving keypoint inputs and sliding perceiver aggregation achieves state-of-the-art translation and alignment on BSL and generalizes to ASL.

Youngjoon Jang, Liliane Momeni, Zifan Jiang, Joon Son Chung and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Paradoxes of Game Theoretic Equilibria and Price of Anarchy

Static equilibrium and black-box regret analysis obscure dynamic disequilibrium; worst-case equilibria are unstable saddles, PoA becomes unbounded under affine costs, and discrete-time learning drives chaos with exponentially degrading inefficiency.

Georgios Piliouras, Ian Gemp, Siqi Liu, Luke Marris

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 1/5
medium 6/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents

PACEvolve++ adapts evolutionary search policies at test time via advisor-model reinforcement learning, using phase-adaptive optimization to outperform frontier-model baselines across engineering and protein tasks.

Minghao Yan, Bo Peng, Benjamin Coleman, Ziqi Chen and 10 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Agent Security is a Systems Problem

Agent security requires systems-level invariants treating AI as untrusted, since model robustness alone cannot prevent real-world agent attacks.

Mihai Christodorescu, Earlence Fernandes, Ashish Hooda, Somesh Jha and 10 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Exploring MLLM-Diffusion Information Transfer with MetaCanvas

MetaCanvas enables multimodal LLMs to plan directly in diffusion latent spaces, outperforming global-conditioning baselines across six precise visual generation tasks.

Han Lin, Xichen Pan, Ziqi Huang, Ji Hou and 9 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 15 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Imperfect World Models are Exploitable

A novel definition of model exploitation reveals it is essentially unavoidable for large policy sets and cannot be precluded in finite ones, yielding safe planning limits.

Logan M Bhamidipaty, Esmeralda S Whitammer, David Abel, Mykel J Kochenderfer and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 2/5
medium 4/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

MIND: Monge Inception Distance for Generative Models Evaluation

MIND uses sliced Wasserstein distance via sorting to evaluate generative models with 10x better sample efficiency, 100x faster computation, and greater adversarial robustness than FID.

Quentin Berthet, Clement CREPY, Romuald Elie, Klaus Greff and 2 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5