Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
Sparse Growing Transformer trains-time sparse depth allocation via progressive attention looping to improve efficiency.
Published 2026Paper ↗

57%
OverallWorth a look
?
OverallWorth a lookVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel1/20reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Progressive attention looping offers a clever training-time sparse depth allocation, though it risks being a benchmark-tuning knob that redistributes rather than reduces compute for an unnamed real-world target.
Abstract
Yao Chen, Yilong Chen, Yinqi Yang, Junyuan Shang, Zhenyu Zhang, Zefeng Zhang, Shuaiyi Nie, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang, Tingwen Liu. Findings of the Association for Computational Linguistics: ACL 2026. 2026.