Good Papers

Showing papers from Carnegie Mellon University / IST-Lisbon (IT) Show all papers

74%Highly rated
?Highly ratedVote to see the score

Sparse Attention as Compact Kernel Regression

Sparse attention corresponds to compact kernel regression, with normalized ReLU and sparsemax arising from Epanechnikov kernels and α-entmax mapping to biweight and triweight kernels. This unifies sparsity with kernel design and yields competitive kernel-based transformers on language modeling and i

Saul Santos, Nuno Gonçalves, Daniel McNamee, Marcos Treviso and 1 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

DashAttention uses adaptive α-entmax to select variable key-value blocks per query, enabling fully differentiable hierarchical sparse attention that matches full-attention accuracy at 75% sparsity with faster inference than FlashAttention-3.

Yuxiang Huang, Nuno Gonçalves, Federico Alvetreti, Lei Li and 4 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5