
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
LLMs distinguish degrees of wrongness among incorrect answers, and alignment with such preferences yields less wrong answers and better calibration.
Published Oct 14, 2024 · 0 citations · Code ★ 10
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.



