
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
Sparse Growing Transformer trains-time sparse depth allocation via progressive attention looping to improve efficiency.
Published 2026 · 0 citations
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.