TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
TETRIS selects optimal draft tokens for batch speculative decoding, improving inference speed and efficiency across varied batch sizes.
Published 2025Paper ↗
70%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel4/20reviewers recommend it
lenient 2/5
medium 2/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Tetris delivers rigorous optimal draft selection for batch speculative decoding, though its real-world latency gains over greedy heuristics remain unverified and it risks being a niche batch-only fix.
Abstract
Zhaoxuan Wu, Zijian Zhou, Arun Verma, Alok Prakash, Daniela Rus, Bryan Kian Hsiang Low. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.