Good Papers

Showing Mixture of experts Show all papers

78%Highly rated
?Highly ratedVote to see the score

Stepped MoE: Segment-Level Routing with Configurable Inference Complexity

Stepped MoE unifies elastic architectures and sparse gating to adapt model capacity to deployment constraints and input requirements, outperforming dense counterparts by 2-5%.

Arnav Kundu, Zhaoyang Xu, Bairu Hou, Chang Gao and 2 more

Published Oct 5, 2026 · ▲ 1 on Hugging Face

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenario

Shuhuai Li, Jianghao Lin, DongDong Ge, Yinyu Ye

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

vExpert: Virtualizing Expert Storage for Adaptive Load Balancing in Distributed MoE Inference

Wenxun Wang, Xiuhong Li, Yida Wang, Chen Tang and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MoEZip: Routing-Aware KV Cache Compression for Sparse Mixture-of-Experts LLMs

Minsang Kim, Seung Baek

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Capricorn: Highly Efficient and Secure Mixture of Experts Inference Framework

Lushan Song, Xiaojian Liang, Shishuai Du, Jun J Sim and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

CECAR: Cache & Expert Co-Aware Routing Accelerates On-Device Inference of MoE LLMs

Byeongju Kim, Hoonki Lee, Byungjun Kim, Inyul Ra and 3 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Towards Principled Fine-Grained MoE Expert Pruning via Pseudo-Boolean Approximation

Zongfang Liu, Ziheng Cheng, Shengkun Tang, Jinghui Zhang and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing

LightMoE replaces redundant MoE experts with parameter-efficient modules to cut memory use 30, 50% while matching or beating existing compression methods.

Jiawei Hao, Zhiwei Hao, Jianyuan Guo, Li Shen and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs

DES selects sequence-level expert coresets for diffusion MoE LLMs to cut unique activations over 55% and latency up to 38% while retaining 99% accuracy, decoupling memory from parallelism.

Hao Chen, Zhiwen Mo, Royson Lee, Qianzhou Wang and 5 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5