
Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment
We introduce CIA to measure chain-of-thought alignment with internal reasoning, finding low alignment that post-training improves substantially while maintaining accuracy.
Published Sep 30, 2026 · 0 citations · ▲ 2 on Hugging Face · Code ★ 1
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.