Good Papers

Showing papers from Standord Show all papers

67%Highly rated
?Highly ratedVote to see the score

MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenario

Shuhuai Li, Jianghao Lin, DongDong Ge, Yinyu Ye

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
86%Must read
?Must readVote to see the score

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice

Proposing geometry-aware online scheduling via Smallest Volume First improves LLM serving's worst-case competitive ratio from 48 to 3 and reduces latency in vLLM.

Li Kong, Qi Qi, Yinyu Ye, Zijie Zhou

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 3/5
78%Highly rated
?Highly ratedVote to see the score

Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale

BalanceRoute uses online F-score routing to cut DP load imbalance and boost LLM serving throughput at scale.

Tianci Bu, Yuan Lyu, Zixi Chen, Chendong Song and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

A Single-Sample Polylogarithmic Regret Bound for Nonstationary Online Linear Programming

A re-solving algorithm achieves O(log² n) regret in nonstationary online linear programming using just one sample per distribution via dynamic programming and dual methods.

Haoran Xu, Owen Shen, Peter W Glynn, Yinyu Ye and 1 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 2/5