Good Papers

Multilinguality in Hybrid Attention LLMs

Hybrid attention LLMs develop cross-lingual alignment tied to recurrent and full-attention layer ordering, with a spike at the first full-attention layer; distillation shows starting with full attention learns up to 2.5× faster.

Lucas Bandarkar, Junlin Hu, Chenyuan Yang, Mohsen Fayyaz, Nanyun Peng

Published Sep 28, 2026▲ 2 on Hugging FacearXiv ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
A striking cross-lingual alignment spike at the first full-attention layer robustly validates layer ordering as a critical multilingual design factor, though the modest layer-swap scope and unanswered recurrence-survival questions keep it from being a definitive architectural…

Abstract

In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to balance the strengths and limitations of full attention and alternatives based on recurrence. This work presents a first study of how hybrid attention impacts the multilinguality of LLMs. Beyond the impact on long sequences in poorly tokenized languages, our study is motivated by the possibility that the inductive biases of the recurrent state alter linguistic processing. Our interpretability analysis confirms this, showing that cross-lingual representations in hybrid models develop in patterns tied to the ordering of recurrent and full-attention layers. Across diverse models, we notably observe a pronounced spike in cross-lingual alignment around the first full-attention layer. These findings lead us to question the conventional ordering of attention layers. In distillation experiments on multilingual data, all alternative layer orderings outperform the standard throughout training, learning up to 2.5X faster. These stark, replicable results prompt our theory that multilingual models would benefit from starting with a full-attention layer rather than recurrent layers.