
Sparse Attention as Compact Kernel Regression
Sparse attention corresponds to compact kernel regression, with normalized ReLU and sparsemax arising from Epanechnikov kernels and α-entmax mapping to biweight and triweight kernels. This unifies sparsity with kernel design and yields competitive kernel-based transformers on language modeling and i
Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
