Good Papers

Showing papers from Habana Labs Show all papers

57%Worth a look
?Worth a lookVote to see the score

More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding

Noam Elata, Itay Lamprecht, Mikey Shechter, Daniel Ohayon and 2 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Block Sparse Flash Attention

Block Sparse Flash Attention accelerates long-context inference by computing exact similarities to select top-k value blocks, skipping ~50% of computation for up to 1.38x kernel and 1.24x end-to-end speedups with minimal accuracy loss.

Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5