HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Hybrid Linear Attention introduces query-dependent chunk-level routing for Gated DeltaNet, improving long-context benchmarks by up to 5.57 points via adaptive recurrent memory composition.
Published Oct 5, 2026 · ▲ 2 on Hugging Face
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.














