Good Papers

Showing Hallucination & factuality Show all papers

70%Highly rated
?Highly ratedVote to see the score

Tuning-Free Accountable Intervention for LLM Deployment -- A Metacognitive Approach

CLEAR enables LLMs to self-identify and correct errors via concept-specific sparse subnetworks without tuning, improving deployment accountability.

Zhen Wah Tan, Jie Peng, Tianlong Chen, Huan Liu

Published Mar 8, 2024 · 3 citations

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 5 of 20 reviewers recommend it
lenient 4/5
medium 1/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Multi-LLM collaboration detects knowledge gaps to make LLMs abstain from wrong answers instead of hallucinating.

Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding and 2 more

Published 2024 · 37 citations

– ReadersNo votes yet
8/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 21 reviewers recommend it
lenient 5/5
medium 3/11
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Teaching LLMs to Abstain across Languages via Multilingual Feedback

Multilingual feedback teaches LLMs to abstain from answering in low-resource languages and improves cross-lingual abstention without degrading performance.

Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding and 5 more

Published 2024 · 4 citations

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

CHORD: Cross-Model Hallucination Detection via Relational Graph Discrimination

Yongxin Deng, Zhen Fang, Guansong Pang, Sharon Li and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Beyond Uniform Detection: Adaptive Hallucination Detection for RAG Across Response Regimes

Jungwuk Park, Sejong Ryu, Jy-yong Sohn, Jaekyun Moon

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation

Shani Goren, Ido Galil, Ran El-Yaniv

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

When Copying Is Hard: Copy-Constrained Decoding for Exact Span Reproduction

Jinghui Zhang, Lang Gao, Zongfang Liu, Ruihong Zeng and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Feeling of Knowing in Large Language Models

Zichuan Fu, Xian Wu, Jingtong Gao, Wenlin Zhang and 8 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation

Jorma Valjakka, Juhani Kivimäki, Juha Mylläri, Jukka K Nurminen

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Why Are LLMs Confidently Wrong? Correcting Overconfident Errors via Causal Head Intervention

Jing Ren, Bowen Li, Ziqi Xu, Xuechao Yang and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Rethinking in Spikes: Mitigating Hallucinations in MDLMs with Step-Aware Decoding

Zhongxing Xu, Zhonghua Wang, Zhe Qian, Shiyan Su and 8 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Large Language Model Failures from Hallucination to Homogenization Are Different Facets of Miscalibration

Tiancheng Hu, Caiqi Zhang, Dirk Hovy, Nigel Collier

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Shared Truth: Emergent Truth Properties in Large Language Models via Heterogeneous Injection-based Transfer

Isaiah Freeman, Joed Ngangmeni

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

PLLS-CP: Unsupervised Hallucination Detection with Layer Selection and Cross-Domain Conformal Guarantees

Vladislav Tsvelenev, Maxim Ryndin

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Alleviating Hallucination with Training-Free Uncertainty-Guided Steering

Litian Liu, Yubing Jian, Qiqi Hou, Reza Pourreza and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SOPO: Socratic Guided Policy Optimization for Span-Level Hallucination Detection

Jiawei Dong, Zilong Bai, Pengtian Zhu, Peng Zhou

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Overthinking as a Symptom of Knowledge Conflict: Understanding and Detecting LLM's Hallucinations in Retrieval-Augmented Question Answering

Zhihua Wen, Li Hao, Zhiliang Tian, Zhizhao Liu and 3 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Invisible Ink, Visible Lies: How Production Watermarking Causes LLMs to Hallucinate

Haocheng Ye, Aoting Hu, Xinwei Zhang, Xunzhu Tang and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Geometry-Calibrated Conformal Abstention for Language Models

Conformal Abstention uses geometry-calibrated confidence to decide abstention with finite-sample correctness guarantees, achieving 75% conditional correctness.

Rui Xu, Yi Chen, Sihong Xie, Hui Xiong

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
88%Must read
?Must readVote to see the score

BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation

BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.

Ning Li, Zixuan Guo, Yan Xu, Wenbo Fei and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
Show 20 more papers