Good Papers

Showing LLM pretraining & scaling laws Show all papers

83%Must read
?Must readVote to see the score

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Looped language models use fixed-point convergence to enable truncated training, shared KV caches, faster prefill, and faster RL updates, while a learned depth prior and orthogonal input injection improve perplexity across scales.

Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen and 3 more

Published Oct 5, 2026 · ▲ 17 on Hugging Face · Code ★ 30

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 2/5
67%Highly rated
?Highly ratedVote to see the score

The Numerical Linear Algebra of Large Language Models

This survey explains large language model core concepts to numerical analysts and highlights key numerical linear algebra contributions to LLM techniques.

Abdelkader Baggag, Yousef Saad

Published Oct 3, 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale

Fixed token codes can replace trainable input embeddings in 1.7B-scale language models, removing 100.7M parameters while preserving substantial capabilities without requiring token-specific vectors.

A. Bochkov

Published Oct 2, 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

Cross-lingual MoE router alignment improves multilingual LLM performance by aligning router outputs across languages instead of hidden states.

Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 1 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
83%Must read
?Must readVote to see the score

Scaling and Distilling Text Embeddings for Better Diffusibility

Scaling and distilling text embeddings improves latent diffusion by yielding more connected, diffusible spaces that boost generative performance beyond autoregressive baselines.

Zekai Zhang, Yunjie Tian, Yanjin He, Xiaoyan Zhang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 57 on Hugging Face · Code ★ 3

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
83%Must read
?Must readVote to see the score

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

LOOM stabilizes looped MoE via bounded residual updates and per-loop routers to scale loops to 9, 12, cutting perplexity from 9.62 to 7.77.

Di He, Pengxiang Li, Da Chang, Qingyan Meng and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 16 on Hugging Face · Code ★ 8

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
80%Must read
?Must readVote to see the score

Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models

Clock Diffusion introduces semi-autoregressive continuous diffusion language models with position-dependent noise schedules, efficient training and sampling, and Cache Grab acceleration to achieve state-of-the-art diffusion likelihoods and competitive reasoning performance.

Yair Schiff, Omer Belhasin, Roy Uziel, Matan Rusanovsky and 6 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

E-MoE improves few-step non-factorized diffusion language models via Mixture-of-Experts routing as a discrete shared latent, boosting sample quality without extra active parameters.

Arseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 66 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 1/5
80%Must read
?Must readVote to see the score

How to Loop MoE: Flatten the Experts, Untie the Attention

Foil improves looped MoE by flattening experts and untying attention, reducing pretraining loss by 0.012 nat and improving routing balance and confidence.

Shouren Wang, Chuang Ma, Mohsen Hariri, Debargha Ganguly and 5 more

Published Sep 28, 2026 · 0 citations · ▲ 9 on Hugging Face · Code ★ 2

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Kimi K3: Open Frontier Intelligence

Kimi K3 is a 2.8 trillion-parameter Mixture-of-Experts model with native vision and 1-million-token context that achieves frontier performance across reasoning, coding, and agentic tasks and outperforms comparable open and proprietary models.

Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao and 36 more

Published Jul 27, 2026 · 0 citations · ▲ 525 on Hugging Face · Code ★ 8,900

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 1/5
83%Must read
?Must readVote to see the score

Scaling Participation in Modular AI Systems

Modular participatory AI combines small stakeholder-trained models into compositional systems that outperform monolithic LLMs by up to 15.4% and exhibit emergent collaborative capabilities.

Shangbin Feng, Yike Wang, Weijia Shi, Luke Zettlemoyer and 2 more

Published Jun 5, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
86%Must read
?Must readVote to see the score

HRM-Text: Efficient Pretraining Beyond Scaling

HRM-Text replaces Transformers with a hierarchical recurrent model and trains on instruction pairs to achieve competitive 1B-parameter performance with 100, 900x fewer tokens and far less compute.

Guan Wang, Changling Liu, Chenyu Wang, Cai Zhou and 5 more

Published May 20, 2026 · 0 citations · ▲ 322 on Hugging Face · Code ★ 2,134

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models

DataFlex unifies sample selection, mixture adjustment, and reweighting for LLMs via a modular LLaMA-Factory framework that improves MMLU and perplexity with faster runtimes.

Hao Liang, Zhengyang Zhao, Mingrui Chen, Meiyi Qiang and 21 more

Published Mar 27, 2026 · 0 citations · ▲ 279 on Hugging Face · Code ★ 2,958

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
80%Must read
?Must readVote to see the score

Attention Residuals

Attention Residuals replace fixed residual accumulation with softmax attention over previous layer outputs for selective, input-dependent aggregation, improving scaling and downstream performance with minimal overhead.

Kimi Team, Guangyu Chen, Yu Zhang, Jianlin Su and 33 more

Published Mar 16, 2026 · 0 citations · ▲ 198 on Hugging Face · Code ★ 3,515

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
89%Must read
?Must readVote to see the score

OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

OPUS defines optimizer-induced update-space data utility for dynamic LLM pre-training selection, outperforming full-scale baselines with minimal overhead.

Shaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu and 8 more

Published Feb 5, 2026 · 0 citations · ▲ 354 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
67%Highly rated
?Highly ratedVote to see the score

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM𝛥 Integration into Upcycled MoE

The method expands multilingual LLMs via post-training PARAMΔ integration into upcycled MoE for data-efficient language acquisition.

Hao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She and 5 more

Published 2026 · 0 citations

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Beyond the Permutation Symmetry of Transformers: The Role of Rotation for Model Fusion

Rotation symmetry generalizes permutation symmetry continuously for transformers, improving parameter matching and model fusion across language and vision tasks.

Binchi Zhang, Zaiyi Zheng, Zhengzhang Chen, Jundong Li

Published Feb 1, 2025 · 0 citations

– ReadersNo votes yet. 2 from authors or colleagues not counted
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Uncovering Scaling Laws for Large Language Models via Inverse Problems

Inverse problem methods uncover scaling laws for large language models, revealing predictive relationships between model size, data, and performance from abstract evidence.

Arun Verma, Zhaoxuan Wu, Zijian Zhou, Xiaoqiang Lin and 14 more

Published 2025 · 0 citations

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Position Paper: Data-Centric AI in the Age of Large Language Models

This position paper argues for data-centric AI in the age of large language models and proposes research directions.

Xinyi Xu, Zhaoxuan Wu, Rui Qiao, Arun Verma and 15 more

Published 2024 · 2 citations

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
80%Must read
?Must readVote to see the score

LMBot: Distilling Graph Knowledge into Language Model for Graph-less Deployment in Twitter Bot Detection

LMBot distills graph neural network knowledge into language models for efficient graph-less Twitter bot detection, achieving state-of-the-art results across four benchmarks.

Cai, Zijian, Zhaoxuan Tan, Zhenyu Lei, Zhu, Zifeng and 3 more

Published Jun 30, 2023 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
89%Must read
?Must readVote to see the score

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

BERT pre-trains deep bidirectional Transformers using masked language modeling and next sentence prediction, achieving state-of-the-art results across language understanding tasks via fine-tuning.

Jacob Devlin, Ming‐Wei Chang, Kenton Lee, Kristina Toutanova

Published 2019 · 34,030 citations

– ReadersNo votes yet
16/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 21 reviewers recommend it
lenient 4/5
medium 9/11
strict 3/5
45%Niche pick
?Niche pickVote to see the score

Rethinking Language Model Scaling under Transferable Hypersphere Optimization

Liliang Ren, Yang Liu, yelong shen, Weizhu Chen

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Meta-memorization and memorization scaling laws in transformers

Alex Nguyen, Kenneth Norman, Gautam Reddy

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Express Language Modeling

Albert Gong, Annabelle M Carrell, Raaz Dwivedi, Lester Mackey

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Position Without Positional Embeddings: A Directional Mechanism in NoPE Transformers

Matan Avitan, Ido Nachum, Yanai Elazar, Yoav Goldberg

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Towards Closing the Autoregressive Gap in Language Modeling via Entropy-Gated Continuous Bitstream Diffusion

Georgios Batzolis, Mark Girolami, Luca Ambrogioni

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Closing the Loop with Fixed-Point Self-Attention

Mrinal Mathur, Barak Pearlmutter, Sergey Plis

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 13 reviewers recommend it
lenient 0/5
medium 0/8
57%Worth a look
?Worth a lookVote to see the score

On the Origin of Algorithmic Progress in AI: Evidence from Language Model Pre-Training

Hans Gundlach, Alex Fogelson, Jayson Lynch, Ana Trisovic and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 16 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/1
57%Worth a look
?Worth a lookVote to see the score

Token-Conditional Expert Dropout: Implicit Regularization for Stable MoE Pretraining

Haili Wang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 16 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/1
45%Niche pick
?Niche pickVote to see the score

Knocking-Heads Attention: Drop-in Shared Projections for Cross-Head Coordination

Zhanchao Zhou, Xiaodong Chen, Haoxing Chen, Zhenzhong Lan and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 14 reviewers recommend it
lenient 0/5
medium 0/9
57%Worth a look
?Worth a lookVote to see the score

Model Collapse is a Singular Complexity Trajectory

Sarwesh Rauniyar

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Pretraining Data Statistics Shape the Phases of Learning Entity Comparison in Language Models

Yik S Chan, Jing Huang, Yanai Elazar, Atticus Geiger

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Self-Distillation of Hidden Layers for Masked Self-Supervised Representation Learning

Scott C. Lowe, Anthony Fuller, Sageev Oore, Evan Shelhamer and 1 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

LIU Hanqing, Jianjun Cao, Yuanze Li, Zijian Zhou

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

On the Information Loss of Multi-Token Prediction: Origin and Solution

Tangyu Jiang, Zhanke Zhou, Haodi Wang, Yiu-ming Cheung and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Geometrically Disentangling Concept Learning from the Language Modeling Loss

Yupei Wang, Neil R Mallinar, Misha Belkin, Alex Warstadt

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

PipeFSDP: Efficient Pipeline Parallel under Fully Sharded Data Parallel for Large Language Model Training

Xinglin Pan, Mingji Han, Penghao Zhao, Lin Zheng and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Towards Precise Knowledge Distillation for Large Language Models via Knowledge Probing

Jiajun Liu, Yao He, Wenjun Ke, Peng Wang and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

PGSB: Pretrained-Guided Shared Basis for LoRA Model Merging

Muqing Liu, Chongjie Si, Zhuoya Liu, Yuheng Jia

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Rethinking Softmax Attention: Polynomial Activations for Transformers

Hemanth Saratchandran, Jianqiao Zheng, Yiping Ji, Wenbo Zhang and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

One Temperature to Rule Them All?

Aviv Orly, Ori Shem Ur, Yaron Oz

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Scaling Laws for SAE Training Data

Nikita Koriagin, Nikita Balagansky, Daniil Gavrilov

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Harder the Task, Sparser the Representation: Sparsity as a Learning Signature of Capability in LLMs

Mingyu Jin, Yutong Yin, Jingcheng Niu, Qingcheng Zeng and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

First-Token Attraction in Mamba Dynamics

Trinh Nguyen, Duy-Tung Pham, Minh-Khoi Nguyen-Nhat, Hoang-Son Do and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

The Cost of Absolute Position: A Spread-Expressivity Tradeoff for Additive Positional Encodings

Noah Mitchell, Isaac Gabriel, Alexander Wyatt

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Circuit-Level Knowledge Distillation for Large Language Models

Yashuo Luo, Tongxu Wang, Siyuan He, Chunyu Wei

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Virtual Head Attention

Xibo Ding, Guoxia Wang, Jinle Zeng, Jiabin Yang and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Inside Emergence: Structure-Behaviour Gaps in Language Model Training

Hak Hyun Kim, Yash Raj, Soroush Vosoughi

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Decomposing and Reshaping Scaling Laws through Token Learning Times

Pingjie Wang, zechenhu, Peiru Yang, Jingtao Han and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Contribution-Aware Structured Sparsity for Model Merging

Yan Li, GUIPING CAO, Meng XU, Tao Jiang and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Speculative Self-Distillation enables Efficient Knowledge Internalization

Shayan Talaei, Agam Bhatia, Arshia Soltani Moakhar, Jonas Hübotter and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SAGE: Semantically Disentangled Representation Learning through Latent Geometry Constraint and Large Language Model

Qiuyu Chen, Liang Xu, Yunnan Wang, Mingqi Yuan and 8 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A Simple Unigram Cross-Entropy Lens on the Lexical Imprint of Pre-training Data

Jeonghoon Kim, Woojin Chung, Woomin Song, Cheonbok Park and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DepthGraft: Structural Regularization through Hierarchical Cross-Layer KV Reconstruction

Xiaohan Qin, Xiangdong Zhang, Yu Wang, Huaijin Wu and 3 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ExpertNavigator: Functionally Coherent Expert Grouping and Pairwise-Ranked Routing for High-Fidelity Dense-to-MoE Conversion

Xiaohan Qin, Cancheng Zhang, Xiaoxing Wang, Xiangdong Zhang and 3 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Why Routers Freeze: Infinite Width Learning Dynamics for Mixture of Experts

Anish Dhir, Volkan Cevher, Leena Chennuru Vankadara

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
Show 20 more papers