Good Papers

Showing Long-context modeling Show all papers

86%Must read
?Must readVote to see the score

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

Hybrid Linear Attention introduces query-dependent chunk-level routing for Gated DeltaNet, improving long-context benchmarks by up to 5.57 points via adaptive recurrent memory composition.

Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He and 2 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
88%Must read
?Must readVote to see the score

Periscope: Extending Frozen Language Models Beyond Their Context Window

Periscope arranges text chunks in a grid to build an evidence map via local and strided probes, letting frozen language models answer questions across multi-million-token contexts with sublinear cost and small GPU memory.

Mohamed Eltahir, Anas Obayd, Raed Rashid, Abdulrahman Alghamdi and 4 more

Published Oct 2, 2026 · ▲ 8 on Hugging Face · Code ★ 1

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
88%Must read
?Must readVote to see the score

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Pretrained transformers stop following references after 1.4, 3.6 lines, but a rank-8 LoRA at one early layer extends computation to 50, 160 lines without changing frozen weights.

Zehao Jin, Ruixuan Deng, 君然 王

Published Sep 29, 2026 · 0 citations · ▲ 75 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling

Triadic linear attention uses 3D tensor states via triadic outer products to scale recurrent state size efficiently, substantially improving long-context modeling and recall.

Oliver Sieberling, Bharat Runwal, David Jin, Ryan Chin and 2 more

Published Sep 29, 2026 · 0 citations · ▲ 27 on Hugging Face · Code ★ 10

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Context Language Models

Context language models treat context as self-modified files to learn context management, outperforming external strategies with lower compute and enabling in-context and parametric learning of management strategies.

Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li and 9 more

Published Sep 29, 2026 · 0 citations · ▲ 43 on Hugging Face · Code ★ 595

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Multilinguality in Hybrid Attention LLMs

Hybrid attention LLMs develop cross-lingual alignment tied to recurrent and full-attention layer ordering, with a spike at the first full-attention layer; distillation shows starting with full attention learns up to 2.5× faster.

Lucas Bandarkar, Junlin Hu, Chenyuan Yang, Mohsen Fayyaz and 1 more

Published Sep 28, 2026 · 0 citations · ▲ 2 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

LongReward: Improving Long-context Large Language Models with AI Feedback

LongReward improves long-context LLMs by applying AI feedback to long-text instruction data via a multi-granularity reward model that evaluates both global coherence and local accuracy.

Jiajie Zhang, Zhongni Hou, Xin Lv, Shulin Cao and 6 more

Published 2025 · 4 citations

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 4 of 20 reviewers recommend it
lenient 2/5
medium 2/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Universal and Efficient Computation with 2D Attention

Christos Tzamos, Guoqing Zheng, Athul Jacob

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Scaling Limits of Long-Context Transformers

Giuseppe Bruno, Chen, Zhengjiang Lin, Yury Polyanskiy and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Short-Context Dominance: How Much Local Context Natural Language Actually Needs?

Vala Vakilian, Zimeng Wang, Ankit Rawat, Christos Thrampoulidis

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Breaking the Exactness Barrier: Interleaved DeepSeek Sparse Attention for Efficient Long Context Reasoning

Yifan GUO, Wei Cui

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Diffusion Language Models Can Approximate Optimal Infilling Lengths Implicitly

Hengchang Liu, Zhao Yang, Bing Su

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AsdaKV: Attention-Overlap Driven Semantic KV Retrieval for Long-Context LLMs

Tianming Yan, Kun Yang, Tingwang You, Kui Ren

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Raven: High-Recall Sequence Modeling via Sparse Memory Routing

Arshia Afzal, Aviv Bick, Eric Xing, Volkan Cevher and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

C3: Long-Horizon Character Consistency via Causal-Continuous State Dynamics and Memory Rewriting

Minghao Chen

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Necessary but Not Sufficient: Spectral Tests for Value-Linear Attention Surrogates

Blaž Škrlj, Ivan Can Arisoy, Solal Vernier

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Context as Low-Rank Weights: Bounded Parametric Dynamic Memory for Unbounded Context

Xiangyu Zhang, Yu Zhou, Taolue Chen

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

State-Resolving Attention for Length-Extrapolating Transformers

Zicheng Liu, Di Wu, Jintao Chen, Di Huang

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CutAttn: Discovering Cognitive Transition Layers for Efficient Long-Context Prefilling

Wentao Liu, Xiabao Wu, Yongchao Liu, Haitao Zhang and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

The Missing Positional Story in LLMs: A Case Study of Shift-Invariant Attention

Benjamin Huh, Hak Hyun Kim, Yuting Tian, Jason Peng and 2 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

GEAR: Generator-Adaptive State Space Models for Associative Recall

Jun Meng, Mohammadhossein Amouei, Zengyu Lin, Xinyu Hu and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

TrajLift: Encoding Verbal Memory Dynamics via Heat Diffusion on Semantic Hierarchies

Jiawen Kang, Jinchao Li, Mingyu Cui, Junan Li and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DeltaFugue: Orchestrating Spatial and Associative Memory for Algorithmic Length Generalization

Anh T Nguyen, Quan Dao, Bing Liu

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Can 4D Foundation Models Remember?

Guangzhao He, Hadar Averbuch-Elor, Wei-Chiu Ma

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Resilient Latent Readouts for Long-Context Question Answering

Jingyi Liao, Wenhao Sun, YITING LI, Zhao Jin and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Prefix Likelihood-Ratio Control: Tail-Stable Training for Long-Horizon Language Generation

Yangyang Liu

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

RoPE Is Not a Proper Relative Position Embedding

Yui Oka, Kyosuke Nishida, Sho Yokoi

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

Tree Rotary Positional Encoding for Extreme Length Extrapolation from Scratch

qiu wu chen, Ziteng Huang, Shuhai Zhang, Zimo Liu and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CHARM+: Cross-Hardware Attention with Re-Merge Multistream Mechanism

Xinhang Zhang, boning zhang, Chengchun Liu, Chunpu Li and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Low-Rank Hierarchical Merging for Efficient Long-to-Short Reasoning

Zeqiu Yu, Xuesheng Zhang, Wenxiao Zhao, Shixiao Wang

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Long-Context Language Models Require Extreme Sparsity in Context Dimension

Prithvi Dixit, Sahil Joshi, Agniva Chowdhury, Anshumali Shrivastava and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 1/5
57%Worth a look
?Worth a lookVote to see the score

AdaCal: Adaptive Calibration for Robust Sparse Attention in Long-Context LLMs

Leqin Xiang, Yongbin Liu, Chunping Ouyang, Ying Yu

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

Siddharth Gollapudi, Prasann Singhal, Nilesh Gupta, Sewon Min

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
83%Must read
?Must readVote to see the score

Language Model Memory and Memory Models for Language

Language model embeddings retain little input information, but autoencoders achieve near-perfect memory; combined causal and retention objectives enable rich, decodable memory formation for efficient encoder-decoder architectures.

Benjamin L Badger

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
89%Must read
?Must readVote to see the score

Dual Dimensionality for Local and Global Attention

Distance-Adaptive Representation uses high-dimensional local and low-dimensional distant keys and values to cut KV cache size while matching full-dimensional baseline performance.

Zhiyuan Wang, Xuan Luo, Sirui Zeng, Xifeng Yan

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
86%Must read
?Must readVote to see the score

When Less is More: The LLM Scaling Paradox in Context Compression

Under lossy context compression, larger compressors reduce reconstruction error but increase unfaithfulness via knowledge overwriting and semantic drift, violating scaling laws for faithful preservation.

Ruishan Guo, Yibing Liu, Guoxin Ma, Yan Wang and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

MemDLM: Memory-Enhanced DLM Training

MemDLM augments diffusion language model training via bi-level optimization with parametric memory, improving convergence, long-context representations, and needle retrieval.

Zehua Pei, Hui-Ling Zhen, Weizhe Lin, Sinno Pan and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 3 on Hugging Face · Code ★ 11

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
83%Must read
?Must readVote to see the score

Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

Gated DeltaNet-2 decouples erase and write with channel-wise gates to improve linear attention, achieving top performance among recurrent and hybrid models at 1.3B scale with strong long-context retrieval.

Ali Hatamizadeh, Yejin Choi, Jan Kautz

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 30 on Hugging Face · Code ★ 326

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

StateTree: Enhancing Long-term Dialogue Reasoning via Reinforcement Learning

StateTree uses RL on tree-structured multi-session path-tracing tasks to train long-context dialogue reasoning, achieving up to +23.60% gains on LongMemEval and outperforming larger baselines while preserving short-context reasoning.

Naen Xu, Wanqing Cui, Yibo Hu, Shixin Hong and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Learning What to Remember: Test-Time Training via Context Distillation

TTCD uses a long-window teacher to supervise a short-window student's fast weights via hidden-state discrepancy, allocating limited memory to future-relevant context and outperforming existing long-context methods with minimal architectural changes.

Zixuan Wang, Xingyu Dang, Rui-Jie Zhu, Zixin Wen and 3 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
83%Must read
?Must readVote to see the score

Efficient Scaling of LLM Training with Flexible Context Parallelism

Flexible Context Parallelism adaptively reconfigures communication groups to eliminate load imbalance and redundant communication, achieving up to 1.46x training speedup over Megatron-LM and DeepSpeed.

Yifan Niu, Han Xiao, Dongyi Liu, Wei zhou and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

LinearARD: Linear-Memory Attention Distillation for RoPE Restoration

LinearARD restores RoPE-scaled LLMs via linear-memory attention self-distillation, recovering 98.3% short-text performance with 4.25M tokens versus 256M.

Ning Yang, Hengyu Zhong, Wentao Wang, Baoliang Tian and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 0/5
89%Must read
?Must readVote to see the score

Fix the Structural Bottleneck: Context Compression via Explicit Information Transmission

ComprExIT fixes structural bottlenecks in LLM context compression via explicit cross-layer feature selection and coordinated transport, improving F1 up to 18.5% with minimal parameters and 2x faster compression.

Jiangnan Ye, Hanqi Yan, Zhenyi Shen, Heng Chang and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 16 on Hugging Face · Code ★ 10

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Learning Evidence Highlighting for Frozen LLMs

HiLight trains a lightweight actor via reinforcement learning to insert highlight tags around pivotal evidence spans in frozen LLM contexts, boosting reasoning without altering inputs or requiring evidence labels.

Shaoang Li, Yanhang Shi, Yufei Li, Mingfu Liang and 9 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
86%Must read
?Must readVote to see the score

Training Transformers for KV-Cache Compressibility

KV-compressibility is a learnable property, so KV-CAT trains transformers via masked KV slots to yield representations more amenable to post-hoc compression without sacrificing quality.

Yoav Gelberg, Yam Eitan, Michael Bronstein, Yarin Gal and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 0/5
91%Must read
?Must readVote to see the score

MAGE: All-[MASK] Block Already Knows Where to Look in Block Diffusion LLM

MAGE uses block-diffusion's aligned all-[MASK] queries to select reusable sparse KV subsets, achieving near-lossless accuracy with up to 6.82x speedup at 128K context.

Omin Kwon, Yeonjae Kim, Doyeon Kim, Minseo Kim and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
78%Highly rated
?Highly ratedVote to see the score

Attention Sinks and Outliers in Attention Residuals

OASIS stabilizes dual-normalized attention-residual architectures via null routing and token-to-depth null coupling, reducing activation outliers by 81.75% and improving low-bit quantized reasoning by 42.11%.

Haozheng Luo, Haoran Dai, Shaoyang Zhang, Xi Chen and 9 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 3

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 1/5
medium 8/10
strict 2/5
86%Must read
?Must readVote to see the score

STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs

STILL introduces self-saliency token selection and norm-preserved feature maps to linearize LLMs, matching original performance with up to 86.2% long-context gains.

Weikang Meng, Liangyu Huo, Yadan Luo, Jiawen Guan and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
91%Must read
?Must readVote to see the score

M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling

M²RNN introduces matrix-valued non-linear RNNs that scale via state expansion, achieving perfect state tracking and outperforming hybrid models with smaller states.

Mayank Mishra, Shawn Tan, Ion Stoica, Joseph Gonzalez and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
89%Must read
?Must readVote to see the score

Multi-Head Recurrent Memory Agents

Multi-Head Recurrent Memory partitions recurrent agent memory into independent heads to prevent overwriting, boosting long-context retention from under 30% to 74% at 896K tokens.

Jiatong Li, Samuel (Min-Hsuan) Yeh, Sharon Li

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

DashAttention uses adaptive α-entmax to select variable key-value blocks per query, enabling fully differentiable hierarchical sparse attention that matches full-attention accuracy at 75% sparsity with faster inference than FlashAttention-3.

Yuxiang Huang, Nuno Gonçalves, Federico Alvetreti, Lei Li and 4 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
83%Must read
?Must readVote to see the score

Recursive Language Models

Recursive Language Models let LLMs recursively process long prompts as external environments, handling inputs two orders of magnitude beyond context windows with large quality gains over baselines at comparable cost.

Alex Zhang, Tim Kraska, Omar Khattab

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 100 on Hugging Face · Code ★ 5,673

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 3/5
80%Must read
?Must readVote to see the score

Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

Proteus progressively expands memory capacity during long-context modeling to reduce interference and boost retention, consistently improving state-of-the-art memory architectures.

Reza Bayat, Ali Behrouz, Vahab Mirrokni, Aaron Courville

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
88%Must read
?Must readVote to see the score

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE

Jet-Long dynamically rescales RoPE via bifocal local and long-range windows with an analytic length-aware schedule to extend LLM contexts without tuning, outperforming baselines on RULER, HELMET-RAG, and perplexity while retaining near-FlashAttention-3 throughput.

Haozhan Tang, Zerui Wang, Yuxian Gu, Song Han and 1 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 24 on Hugging Face · Code ★ 13

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5