
Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation
LLM novelty judges are unstable: small prompt changes alter verdicts on over half of identical idea pairs and shift accuracy by over 50 points, undermining automated ideation evaluations.
Published Oct 1, 2026 · 0 citations
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.



