45%Niche pick?Niche pickVote to see the scoreNeurIPS 2026Imperial College LondonEPFLGatsby Unit, University College LLM pretraining & scaling lawsWhy Routers Freeze: Infinite Width Learning Dynamics for Mixture of ExpertsAnish Dhir, Volkan Cevher, Leena Chennuru VankadaraSydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet0/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 0 of 20 reviewers recommend itlenient 0/5medium 0/10strict 0/5
86%Must read?Must readVote to see the scoreNeurIPS 2026Gatsby Unit, University College AmazonUniversity College London, UniveU TübingenLLM pretraining & scaling lawsHow to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable ParameterizationResearchers derive maximally scale-stable parameterizations for Mixture-of-Experts via dynamical mean-field theory, yielding robust learning-rate transfer and monotonic scaling gains across regimes.Leena Chennuru Vankadara, Moritz Haas, Luke Hayward, Sebastian Bordt and 1 moreParis Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026 · ▲ 1 on Hugging Face · Code ★ 4– ReadersNo votes yet14/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 14 of 20 reviewers recommend itlenient 3/5medium 8/10strict 3/5