Good Papers

Transformers Provably Implement In-Context Reinforcement Learning with Policy Improvement

Linear self-attention transformers provably implement in-context policy-improvement via explicit constructions, with gradient flow converging exponentially to optimal RL update parameters under distribution richness conditions.

Haodong Liang, Lifeng LAI

Published 2026Atlanta Poster Session 5 · Fri, Dec 11, 10:00 AM–1:00 PM local time · Hall C1arXiv ↗OpenReview ↗

78%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel11/20reviewers recommend it
lenient 3/5
medium 5/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

We investigate the ability of transformers to perform in-context reinforcement learning (ICRL), where a model must infer and execute learning algorithms from trajectory data without parameter updates. We show that a linear self-attention transformer block can provably implement policy-improvement methods, including semi-gradient SARSA and actor-critic, via explicit parameter constructions. Beyond existence, we design a teacher-mimicking training procedure, analyze its gradient-flow dynamics, and establish the first convergence guarantee in the ICRL literature: under suitable richness conditions on the training MDP distribution, gradient flow converges locally and exponentially to an optimal parameter manifold corresponding to the desired RL update. Empirically, training transformers on randomly generated tabular MDPs confirms these predictions: the learned models recover the parameter structure of our explicit constructions and, when deployed on unseen MDPs, deliver strong in-context control performance. Together, these results illuminate how transformer architectures internalize and execute classical reinforcement learning algorithms in context, bridging mechanistic understanding and training dynamics in ICRL.