A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning
Online autoregressive learning mistake bounds grow from constant to logarithmic in generation horizon M with end-to-end feedback, but chain-of-thought access removes M dependence entirely.
Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
