Good Papers

Showing Efficient attention & state-space models Show all papers

57%Worth a look
?Worth a lookVote to see the score

Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping

Sparse Growing Transformer trains-time sparse depth allocation via progressive attention looping to improve efficiency.

Yao Chen, YiLong Chen, Yinqi Yang, Junyuan Shang and 8 more

Published 2026 · 0 citations

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Nearly Optimal Attention Coresets

Alexandr Andoni, Eldar Kleiner, Edo Liberty

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache

Mohsen Dehghankar, Abolfazl Asudeh

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

Ngoc Bui, Trung Hieu Nguyen, Arman Cohan, Rex Ying

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · Code ★ 21

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Mixture-of-Top-$k$ Attention: Efficient Attention as Scalable Fast Weights

Qishuai Wen, Zhiyuan Huang, meng xianghan, Wei He and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Basis-Mediated Bilinear Attention: A New Method for Greatly Reducing Query--Key Pathway Parameters

Wei Chen, Wenhao Jiang, Yiying Yang

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Parallel Broyden methods for efficiently evaluating nonlinear state space models

Ian Christopher Tanoh, Scott Linderman

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding

Noam Elata, Itay Lamprecht, Mikey Shechter, Daniel Ohayon and 2 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

DisSparse: Pipelined Top-$p$ Sparse Attention for Long-Context LLM Serving

Nurlan Nazaraliyev, Elaheh Sadredini, Nael Abu-Ghazaleh

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Prism Attention: Proposal-Refined Index Sharing Mechanism for Efficient LLMs Inference

Zhenxu Tian, Kebin Liu, Zhengwu Yang, Yi Su and 7 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

BitShift-RoPE: Zero-FLOP Relative Positional Encoding for Spiking Neural Network Transformers

Seung-Kyu Hong, Sangheum Hwang, HYUK-YOON KWON

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Bridging the Gap: Position-Independent Cache Reuse for Hybrid SSM-Attention Architectures

Weikang Wang, Xin Zhou, Haiyang Liu, Lei Wang and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

GPU Hierarchy Meets Structured Matrices: Fast Algorithms for State-Space Models

Berlin Chen, Caitlin Wang, Aakash Sunil Lahoti, Kevin Li and 7 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Task-Aware KV Cache Compression for LLM Agents via Utility-Driven Step Pruning

Yusen Wu, Yefan Wang, Jia Yee Tan, Guangyuan Dong and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

LAPrune: Logits-Aligned Scoring Proxy for KV Pruning via Vector Quantization

Mingyang Yu, Rong-Cheng Tu, Yifu Ding, Hanqing Zhao and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State

Kaicheng Xiao, Haotian Li, Liran Dong, Guoliang Xing

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Vortex: Efficient and Programmable Sparse Attention Serving

Zhuoming Chen, Xinrui Zhong, Qilong Feng, Ranajoy Sadhukhan and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

FlashMask-3: Efficient and Expressive Mask-Aware Distributed Attention

Guoxia Wang, Qianyue He, Siming Wu, Haoyang Xie and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SpectralKV: Redundancy-Aware KV Cache Compression via Spectral Coreset Selection

Fangming Zhao, Fulun Ye, Xiaofei Yue, Ziming Zhao and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Token Filtering: Online Attention Pruning via KV Similarity for Efficient LLM Inference

JUNGMIN LEE, Gwangeun Byeon, Yulhwa Kim, Seokin Hong

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

The Kernel Reality Check: Benchmarking and Distilling Efficient Attention at Scale

Firat Oncel, Cem Subakan, Mirco Ravanelli, Çağatay Yıldız

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling

HeatKV ranks VAR attention heads by cross-scale attention to build static pruning schedules, doubling KV-cache compression versus prior methods while preserving image quality on Infinity-2B.

Jonathan Cederlund, Axel Berg, Durmus Alp Emre Acar, Chuteng Zhou and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases

Positional LSH represents ALiBi's bias matrix via binary block masks, yielding near-linear approximate attention with uniform accuracy across inputs.

Daniel Wolfson, Tal Wagner

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection

CompactAttention accelerates chunked prefill via block-union KV selection that builds minimal per-group block tables to eliminate KV copy overhead, achieving up to 2.72x attention speedup with near-dense accuracy on 128K contexts.

Jiwon Song, Dongwon Jo, Beomseok Kang, jae-joon kim

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 12 on Hugging Face · Code ★ 7

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 0/5
80%Must read
?Must readVote to see the score

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

CommunityKV formulates sparse attention as graph community detection to retrieve coherent token groups via constant-time updates, achieving up to 1.71x long-context decoding throughput.

Joe McKenna, Anastasios Alexandridis, Nathan Susanj, Jing Liu

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Approaching I/O-optimality for Approximate Attention

Approximate attention algorithms achieve near-linear I/O cost in sequence length via efficient approximate methods with matching lower bounds proving near-optimality.

Pál András Papp, Aleksandros Sobczyk, Anastasios Zouzias

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 3/5
71%Highly rated
?Highly ratedVote to see the score

Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers

Systematic analysis of triangular inversion for delta-rule linear transformers yields algorithms with up to 4.3x speedup on NPUs and preserved end-to-end accuracy across low-precision settings.

Aleksandros Sobczyk, Gioele Gottardo, Christos K Matzoros, Mirko De Vita and 3 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
88%Must read
?Must readVote to see the score

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

RTPurbo converts full-attention LLMs into sparse models within hundreds of steps via retrieval heads and dynamic indexing, achieving near-lossless accuracy with 9.36x prefill and 2.01x decode speedups.

Yanke Zhou, Yiduo Li, Hanlin Tang, Maohua Li and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 90 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Block Sparse Flash Attention

Block Sparse Flash Attention accelerates long-context inference by computing exact similarities to select top-k value blocks, skipping ~50% of computation for up to 1.38x kernel and 1.24x end-to-end speedups with minimal accuracy loss.

Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Flexformer: Flexible Linear Transformer with Learnable Attention Kernel

Flexformer learns attention kernels via trainable spectral frequencies in linear attention, outperforming baselines on long-sequence tasks and enabling efficient Transformer distillation.

Haoran Zhang, Zisu Dong, Feng Zhou

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
80%Must read
?Must readVote to see the score

Benchmarking Attention for Tabular Foundation Models

This paper benchmarks 2D tabular attention across GPU backends, finding optimal choices vary by row versus column attention, hardware, and sequence length.

Maximilian Schambach, Clemens Biehl, Sam Thelin

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 3/5
86%Must read
?Must readVote to see the score

The Key to Going Linear: Analysis-Driven Transformer Linearization

Analysis-driven transformer linearization isolates state update design to show delta-style networks outperform gated accumulation via key-dependent rank-1 projections, reducing approximation errors with sink tokens and cache routing to match adaptive caching at 32B scale.

Anna Kuzina, Paul Whatmough, Babak Ehteshami Bejnordi

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference

Flux Attention dynamically routes layer-level attention between full and sparse modes via a lightweight router to accelerate LLM inference, achieving up to 2.8x prefill and 2.0x decode speedups with minimal training.

Quantong Qiu, Zhiyi Hong, Yi Yang, Haitian Wang and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5