Good Papers

NeurIPS 2026 posters

Best rated first.

93%Must read
?Must readVote to see the score

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

RL post-training yields progress advantage, a log-ratio that recovers optimal step-level advantage without dedicated reward models, outperforming trained alternatives across agent benchmarks.

Changdae Oh, Wendi Li, Seongheon Park, Samuel (Min-Hsuan) Yeh and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 13 on Hugging Face · Code ★ 12

100% Readers1 of 1 upvoted
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
92%Must read
?Must readVote to see the score

SpatialBench: Is Your Spatial Foundation Model an All-Round Player

SpatialBench evaluates 41 spatial foundation models across 19 datasets and finds none are all-round players, with full-context attention maximizing accuracy and domain alignment exceeding scaling for embodied tasks, plus it introduces DA-Next-5M and DA-Next.

Haosong Peng, Hao Li, jiaqi chen, Yuhao Pan and 9 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 69 on Hugging Face · Code ★ 138

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
92%Must read
?Must readVote to see the score

Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition

Root cause analysis benchmarks conflate retrieval and reranking failures, revealing graph methods rarely beat statistical baselines; a two-stage retriever-LLM reranker matches or exceeds all baselines without causal graphs or labels.

Hada M Muhammad, Luan Pham, Laure Barrière, Sachin Shetty and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
92%Must read
?Must readVote to see the score

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

K12-KGraph introduces a curriculum-aligned K-12 knowledge graph, benchmark, and training data showing current LLMs achieve under 57 percent accuracy on curriculum cognition and that graph-guided supervision outperforms generic instruction tuning.

Hao Liang, Qihan Lin, Mingrui Chen, Hengyi Feng and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 62 on Hugging Face · Code ★ 392

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
92%Must read
?Must readVote to see the score

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

MCP-Atlas benchmarks LLM tool-use on 1,000 real-server tasks, finding frontier models reach 82.2% pass rates but 63.3% of failures are cognitive.

Chaithanya Bandi, Razvan Dumitru, Ben Hertzberg, Divyansh Agarwal and 15 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 5/5
92%Must read
?Must readVote to see the score

OdysSim: Building Foundation Models for Human Behavior Simulation

OdysSim trains 8B behavioral foundation models via SOUL taxonomy and multi-stage recipes, ranking first on eight human simulation benchmarks while nearly matching real-user reaction alignment.

Xuhui Zhou, Weiwei Sun, Weihua Du, Jiarui Liu and 5 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
92%Must read
?Must readVote to see the score

SR-Prominence: A Crowdsourced Protocol and Dataset Suite for Perceptually-Weighted Super-Resolution Artifact Evaluation

SR-Prominence defines artifact prominence via crowdsourced annotations across 3,935 masks and shows classical full-reference metrics surprisingly detect perceptual impact better than specialized detectors.

Ivan Molodetskikh, Kirill Malyshev, Mark Mirgaleev, Nikita Zagainov and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 5/5
92%Must read
?Must readVote to see the score

CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving

CausalDriveBench evaluates causal reasoning in autonomous driving vision-language-action models via structured QA and counterfactual trajectories, finding weak causal understanding despite fluent reasoning and accurate baseline predictions.

Narendiran Chembu, Navvrat Rao, Shreedhar Kodate, Gayatri S Banda and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 5/5
92%Must read
?Must readVote to see the score

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

Low-bit KV cache quantization silently collapses LLM safety alignment via geometric subspace vulnerability, and per-channel reduction diagnostics recover up to 97% of lost refusals.

Bruce C Xu, Adarsh Kumarappan, Mu Zhou

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
92%Must read
?Must readVote to see the score

LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

LowRankArena standardizes SVD-based LLM compression evaluation and reveals that method rankings and speedups depend heavily on backbone and workload under aligned protocols.

Zishan Shao, Lixun Zhang, Kangning Cui, Wenhao Wu and 9 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
92%Must read
?Must readVote to see the score

MemLeak: Diagnosing Information Leaks in Multimodal Agent Memory

Multimodal AI agents retain forgotten facts via implicit visual cues, with MemLeak showing 12% image-based recovery and content-aware deletion reducing residuals to 2%.

Kuan Wang, Chao Zhang

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 5/5
92%Must read
?Must readVote to see the score

MorphoHELM: A Comprehensive Benchmark for Evaluating Representations for Microscopy-Based Morphology Assays

MorphoHELM benchmarks microscopy representation methods across batch effects, finding classic computer vision strategies outperform deep learning across settings and revealing trade-offs between models.

Emre Hayir, Lorin Crawford, Alex X Lu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
92%Must read
?Must readVote to see the score

Reinforcement Learning for Code Optimization

Reinforcement learning for code optimization fails due to noisy, sparse execution-time rewards, so a calibrated three-stage pipeline improves strict pass rates by up to 125% while preserving correctness.

Pierre Chambon, Kunhao Zheng, Juliette Decugis, Benoît Sagot and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 13 on Hugging Face

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 5/5
92%Must read
?Must readVote to see the score

Measuring Cross-Modal Synergy: A Benchmark for VLM Explainability

Cross-modal redundancy causes unimodal metrics to contradict (τ=-0.06), so Synergistic Faithfulness (F_syn) isolates joint modality dividends with ρ=0.92 and 24× speedup, revealing VLM explainers over-index visual salience versus adapted attention methods.

Joël Roman Ky, Salah GHAMIZI, Maxime Cordy

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
92%Must read
?Must readVote to see the score

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

TRL-Bench standardizes cross-paradigm evaluation of tabular encoders via shared representation-level probes, finding encoder quality is task-specific and best pipelines combine capability-matched specialists.

Wei Pang, Xiangru Jian, Hehan Li, Zhixuan Yu and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 54 on Hugging Face · Code ★ 10

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
92%Must read
?Must readVote to see the score

DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding

DocScope benchmarks verifiable long-document reasoning via structured trajectory evaluation, finding correct answers rarely include complete evidence chains and region grounding remains weakest.

Xiang Feng, Jiawei Zhou, Zhangfeng Huang, Kewei Wang and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 5/5
92%Must read
?Must readVote to see the score

Log-Likelihood, Simpson’s Paradox, and the Detection of Machine-Generated Text

Average token-level log-likelihood scores suffer Simpson’s paradox across hidden-space regions, and local calibration via learned score-distribution predictors fixes it, boosting detection AUROC substantially.

Tom Kempton, Viktor Drobnyi, Maeve Madigan, Stuart Burrell

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
92%Must read
?Must readVote to see the score

Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking

LLM pipeline evaluation variance is underestimated because design choices are ignored, so corrected intervals restore coverage and cut benchmark gaming.

Solomon Messing

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
91%Must read
?Must readVote to see the score

Attention Transfer Is Not Universally Effective for Vision Transformers

Attention transfer fails for four ViT families due to architectural mismatch, and adding the teacher's native components to students fully restores its effectiveness.

Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Peng Hu and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 4/5
91%Must read
?Must readVote to see the score

Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning

A conformal procedure for chain-of-thought reasoning replaces majority voting with calibrated weighted aggregation to provide finite-sample confident-error guarantees and improves selective accuracy without retraining.

Yu Gu, Zijun Yu, Vahid Partovi Nia, Masoud Asgharian

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

Self Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale

PubMed is autonomously converted into structured biomedical datasets larger, more nuanced, and more accurate than manual repositories via ontology tagging, hybrid retrieval, and a multi-agent extraction system.

Haydn Jones, Yimeng Zeng, Alden Rose, Yifei Li and 10 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
91%Must read
?Must readVote to see the score

DriveSpatial: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving

DriveSpatial benchmarks vision-language models' spatiotemporal autonomous driving intelligence, finding a 28.4-point human gap with cognitive scene construction as the key bottleneck.

Anh Hao Vo, Khoa Vo, Phu Loc Nguyen, Sieu Tran and 9 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

Models Recall What They Violate: Constraint Adherence in Multi-Turn LLM Ideation

DriftBench finds iterative LLM ideation increases complexity and reduces constraint adherence, with models often violating rules they accurately recall and judges under-detecting violations.

Garvin Kruthof

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 5/5
91%Must read
?Must readVote to see the score

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation

VeriContest introduces 946 competitive programming problems with verified Rust specifications and proofs, showing state-of-the-art models reach only 5.29% on end-to-end verifiable generation.

Zichen Xie, Mrigank Pawagi, Yuxin Liu, Aaditi Rai and 4 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
91%Must read
?Must readVote to see the score

Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs

Token Inoculation conditions LLMs to retain dual-use knowledge gated by a special token, reducing hazardous accuracy to 18% while preserving 93% of benign performance across 1B-14B scales.

Seung-Hyun Lee, Dongyoon Han, Sangdoo Yun

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

Beyond LoRA vs. Full Fine-Tuning: Gradient-Guided Optimizer Routing for LLM Adaptation

MoLF dynamically routes optimizer updates between full fine-tuning and LoRA to match or beat the stronger static method across tasks, and its efficient variant surpasses AdaLoRA and AdaMix by up to 11.70 points.

Haozhan Tang, Xiuqi Zhu, Xinyin Zhang, Boxun Li and 2 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

A two-parameter scaling law quantifies diminishing returns in multi-agent LLM systems, showing that dense debate hits hard ceilings, noise placebos match self-correction, and only heterogeneous teams escape diminishing returns.

Blaz Bertalanic, Carolina Fortuna

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 4/5
91%Must read
?Must readVote to see the score

CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations

CardioLens evaluates MLLMs on multi-sequence cardiac MRI, revealing poor clinical workflow performance and category-collapse failures despite reasoning prompts and slice selection.

Zixian Su, Hongkai Zhang, Fan Gao, Encheng Su and 11 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 5/5
91%Must read
?Must readVote to see the score

Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents

Agents encountering benign errors suffer "accidental meltdowns", unsafe behaviors like unauthorized reconnaissance, across 64.7% of error rollouts, often unreported.

Rishi Jha, Harold Triedman, Vitaly Shmatikov, Arkaprabha Bhattacharya

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
91%Must read
?Must readVote to see the score

FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

FineVLA introduces fine-grained action-aligned supervision for steerable vision-language-action policies, yielding up to 86.8% simulation and 62.7 real-world success and boosting steerable control over coarse instructions.

Xintong Hu, Xuhong Huang, JINYU ZHANG, Yutong Yao and 8 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

RealityTest: How People Probe AI Identity and Whether Models Disclose It

RealityTest benchmarks multimodal multilingual AI identity disclosure via 3,152 human queries, finding question phrasing and context dominate over model choice and suppression cuts rates below 30%.

Anna Gausen, Sarenne Wallbridge, Bessie O'Dell, Christopher Summerfield and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
91%Must read
?Must readVote to see the score

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs

StereoTales reveals open-ended LLM generation emits shared harmful stereotypes that culturally adapt to prompt languages and align with human harmfulness ratings.

Pierre Le Jeune, Etienne Duchesne, Weixuan Xiao, Stefano Palminteri and 3 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
91%Must read
?Must readVote to see the score

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Claw-Eval introduces a trajectory-aware benchmark with 300 tasks, finding opaque grading misses 44% of safety violations and agent rankings vary across multi-dimensional capabilities.

Bowen Ye, Rang Li, Qibin Yang, Yuanxin Liu and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 117 on Hugging Face · Code ★ 778

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
91%Must read
?Must readVote to see the score

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech

DUET enables plug-and-play emotion control for pretrained diffusion and flow-matching TTS by steering hidden states and guiding mel-spectra via a differentiable vocoder, surpassing supervised emotional baselines.

Xu Zhang, Longbing Cao, zhangkai wu

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

When Attention Collapses: Residual Evidence Modeling for Compositional Inference

Under additive superposition, attention slots collapse to dominant components because memoryless attention ignores explained evidence; residual evidence depletion prevents collapse and enables compositional inference.

Niklas Houba

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

P$^{3}$: Joint Program-and-Proof Planning\\ for Verified Code Generation

P³ plans programs and proofs jointly from specifications before elaboration, outperforming sequential baselines by up to 11.2 points on verified generation benchmarks while reducing cost and time.

Zenan Li, Ziran Yang, Peiyang Song, Zhaoyu Li and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents

MemPoison benchmarks 1227 adversarial cases across memory substrates and finds write-time defenses fail against multi-record and dormant corruption, requiring adaptive defenses.

Jifeng Gao, Kang Xia, Yi Zhang, Xiaobin Hong and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
91%Must read
?Must readVote to see the score

Jointly Reinforcing Diversity and Quality in Language Model Generations

DARLING uses a learned partition function to jointly optimize language model response quality and semantic diversity via reinforcement learning, improving both quality and novelty across creative and math benchmarks.

Tianjian Li, Yiming Zhang, Ping Yu, Swarnadeep Saha and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 25 on Hugging Face · Code ★ 61

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

Beyond the Golden Teacher: Enhancing Graph Learning through LLM-GNN Co-teaching

LLM-GNN Co-Teaching replaces golden-teacher design with bidirectional pseudo-label exchange and trajectory-based preference optimization, boosting few-shot graph accuracy by up to 7.86%.

Zhuoyi Peng, Hanlin Gu, Lixin Fan, Yi Yang

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
91%Must read
?Must readVote to see the score

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing

Contrastive Decoding Diffing recovers verbatim implanted facts and pipeline artifacts via output-level logit differences without weight access, outperforming white-box methods 170x faster.

Michał Brzozowski, Zuzanna Dubanowska, Enrico Cassano, Neo Christopher Chung

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
Show 40 more papers