Good Papers

Showing Efficient inference & serving Show all papers

78%Highly rated
?Highly ratedVote to see the score

Closing the Context Gap: Activation Alignment for Tabular In-Context Learning

Activation alignment trains a linear map to align partial-context student activations with full-context teacher activations, significantly improving tabular in-context learning efficiency and recovering much of the performance gap.

Yoel Zeldes

Published Oct 5, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
88%Must read
?Must readVote to see the score

Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

Post-training multi-token prediction heads on ~2.5B chain-of-thought tokens match pretraining speedups with 10^3-10^4x less data, while chain-aware verification and adaptive head selection boost throughput up to 16%.

Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne and 7 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
91%Must read
?Must readVote to see the score

Learning Functional Subspaces for Neural Network Compression

LSP learns low-rank subspaces end-to-end via joint orthogonal projector optimization to reduce transformer memory and compute while outperforming local criteria at high compression ratios.

Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 1 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
89%Must read
?Must readVote to see the score

SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving

SlimWise decouples MoE expert pruning across prefill and decode phases to boost serving throughput without sacrificing accuracy via direct KV cache reuse and selective distillation.

Gunho Park, Kyoungho Jeun, Juntaek Oh, Byeongjun Shin and 2 more

Published Sep 28, 2026 · 0 citations · ▲ 5 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
80%Must read
?Must readVote to see the score

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

HeteroFold enables prefill-free cross-family KV cache transfer between frozen heterogeneous LLM agents, accelerating 32K context transfer up to 10.7x while matching text-based multi-agent performance.

Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo and 2 more

Published Sep 26, 2026 · 0 citations · ▲ 95 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
88%Must read
?Must readVote to see the score

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

Edge0 predicts next-layer MoE routing one token ahead to stream experts from SSD, serving 35B-class MoEs at 20 tok/s within 3 GiB active memory on a 24 GB machine via recovery LoRA adapters.

Yu Lin, Yiming Wang, Runyuan Cai, Liu, Hanze and 1 more

Published Sep 16, 2026 · 0 citations · ▲ 25 on Hugging Face · Code ★ 3,293

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
70%Highly rated
?Highly ratedVote to see the score

TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding

TETRIS selects optimal draft tokens for batch speculative decoding, improving inference speed and efficiency across varied batch sizes.

Zhaoxuan Wu, Zijian Zhou, Arun Kumar Verma, Alok Prakash and 2 more

Published 2025 · 0 citations

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 4 of 20 reviewers recommend it
lenient 2/5
medium 2/10
strict 0/5
89%Must read
?Must readVote to see the score

Efficient Memory Management for Large Language Model Serving with PagedAttention

PagedAttention applies OS virtual memory and paging to LLM key-value caches, reducing waste and duplication to boost vLLM throughput 2-4x over existing systems.

Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng and 5 more

Published Sep 12, 2023 · 50 citations · ▲ 76 on Hugging Face · Code ★ 86,094

– ReadersNo votes yet
16/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 21 reviewers recommend it
lenient 5/5
medium 8/11
strict 3/5
45%Niche pick
?Niche pickVote to see the score

DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs

Daewon Chae, Hyunwon Chung, Changwoo Lee, Hun-Seok Kim

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Map-Guided Caching: A Global Perspective for Efficient Diffusion Transformer

Jiaqi Ji, Ran Yang, Bo Wei, Hui Li and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

The $1/\mathcal{W}$ Law: Context Length is the Dominant Energy Lever in LLM Inference Fleets

Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Adaptive Entropy-Sparing for Efficient Reasoning

Cong Jiang, Xiaofeng Zhang, Tom Ko, Zheng Zhang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

TokenRouter: Efficient Serving System for Token-Level LLM Routing

Tianyu Fu, Tengxuan Liu, Ruoxi Wang, Yixin Dong and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

LLMs Optimizing LLMs: Automated MegaKernel Generation for Inference Acceleration

Weiqiang Xiong, Shaohui Peng, Wenyi Li, Hao Lu and 7 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Accelerating the Inference Era with AI-Driven, Globally Optimized HW/SW Co-Design

Miria Feng, Fangzhao Zhang, Adrian G Lafuente, Mert Pilanci and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Efficient Tree Draft for Long-Context Speculative Decoding

Brian J Chan, Ning-Chi Huang, Kai-Chiang Wu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

VALOR: Vector-Aware Low-Rank Restructuring of Neural Networks for RISC-V Inference

Zhihao Xu, Xiaoning Du, Bixin Li, Wang Lulu and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Pacing Branch Parallelism in LLM Serving

Swapnil Gandhi, Siva Kumar Sastry Hari, Bill Dally, Christos Kozyrakis

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SPECS: Faster Test-Time Scaling through Speculative Drafts and Dynamic Switching

Mert Cemri, Nived Rajaraman, Rishabh Tiwari, Xiaoxuan Liu and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

μLM: Rethinking Sub-100M Language Models through Memory-First Design

Zijie Chen, Guiyun Fan, Haiming Jin

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
Show 20 more papers