Sparse attention corresponds to compact kernel regression, with normalized ReLU and sparsemax arising from Epanechnikov kernels and α-entmax mapping to biweight and triweight kernels. This unifies sparsity with kernel design and yields competitive kernel-based transformers on language modeling and i
DashAttention uses adaptive α-entmax to select variable key-value blocks per query, enabling fully differentiable hierarchical sparse attention that matches full-attention accuracy at 75% sparsity with faster inference than FlashAttention-3.