TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
TETRIS selects optimal draft tokens for batch speculative decoding, improving inference speed and efficiency across varied batch sizes.
Published 2025 · 0 citations
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
