Good Papers

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

Reinforcement learning post-training yields decomposed-to-composed asymmetry in language models: decomposed training rarely transfers to composed tasks, while composed training readily transfers backward, with preliminary evidence extending to real-world tool use.

Yu He, Yingxi Li, Yifei Wang, Ellen Vitercik

Published 2026Atlanta Poster Session 5 · Fri, Dec 11, 10:00 AM–1:00 PM local time · Hall C1arXiv ↗OpenReview ↗

83%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel13/20reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation and clear compositional structure. We find a consistent decomposed-to-composed asymmetry: decomposed-skill training does not reliably transfer to composed tasks, whereas composed-task training transfers more readily back to decomposed tasks. We provide theoretical explanation for this asymmetry, and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, we present a pilot study on real-world tool-calling benchmarks, showing preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings.