Good Papers

Showing LLM pretraining & scaling laws Show all papers

83%Must read
?Must readVote to see the score

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Looped language models use fixed-point convergence to enable truncated training, shared KV caches, faster prefill, and faster RL updates, while a learned depth prior and orthogonal input injection improve perplexity across scales.

Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen and 3 more

Published Oct 5, 2026 · ▲ 20 on Hugging Face · Code ★ 35

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 2/5
67%Highly rated
?Highly ratedVote to see the score

The Numerical Linear Algebra of Large Language Models

This survey explains large language model core concepts to numerical analysts and highlights key numerical linear algebra contributions to LLM techniques.

Abdelkader Baggag, Yousef Saad

Published Oct 3, 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale

Fixed token codes can replace trainable input embeddings in 1.7B-scale language models, removing 100.7M parameters while preserving substantial capabilities without requiring token-specific vectors.

A. Bochkov

Published Oct 2, 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

Cross-lingual MoE router alignment improves multilingual LLM performance by aligning router outputs across languages instead of hidden states.

Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 1 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Scaling and Distilling Text Embeddings for Better Diffusibility

Scaling and distilling text embeddings improves latent diffusion by yielding more connected, diffusible spaces that boost generative performance beyond autoregressive baselines.

Zekai Zhang, Yunjie Tian, Yanjin He, Xiaoyan Zhang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 57 on Hugging Face · Code ★ 3

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

LOOM stabilizes looped MoE via bounded residual updates and per-loop routers to scale loops to 9, 12, cutting perplexity from 9.62 to 7.77.

Di He, Pengxiang Li, Da Chang, Qingyan Meng and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 16 on Hugging Face · Code ★ 8

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models

Clock Diffusion introduces semi-autoregressive continuous diffusion language models with position-dependent noise schedules, efficient training and sampling, and Cache Grab acceleration to achieve state-of-the-art diffusion likelihoods and competitive reasoning performance.

Yair Schiff, Omer Belhasin, Roy Uziel, Matan Rusanovsky and 6 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

E-MoE improves few-step non-factorized diffusion language models via Mixture-of-Experts routing as a discrete shared latent, boosting sample quality without extra active parameters.

Arseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 66 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

How to Loop MoE: Flatten the Experts, Untie the Attention

Foil improves looped MoE by flattening experts and untying attention, reducing pretraining loss by 0.012 nat and improving routing balance and confidence.

Shouren Wang, Chuang Ma, Mohsen Hariri, Debargha Ganguly and 5 more

Published Sep 28, 2026 · 0 citations · ▲ 9 on Hugging Face · Code ★ 2

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Kimi K3: Open Frontier Intelligence

Kimi K3 is a 2.8 trillion-parameter Mixture-of-Experts model with native vision and 1-million-token context that achieves frontier performance across reasoning, coding, and agentic tasks and outperforms comparable open and proprietary models.

Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao and 36 more

Published Jul 27, 2026 · 0 citations · ▲ 526 on Hugging Face · Code ★ 8,899

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Scaling Participation in Modular AI Systems

Modular participatory AI combines small stakeholder-trained models into compositional systems that outperform monolithic LLMs by up to 15.4% and exhibit emergent collaborative capabilities.

Shangbin Feng, Yike Wang, Weijia Shi, Luke Zettlemoyer and 2 more

Published Jun 5, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

HRM-Text: Efficient Pretraining Beyond Scaling

HRM-Text replaces Transformers with a hierarchical recurrent model and trains on instruction pairs to achieve competitive 1B-parameter performance with 100, 900x fewer tokens and far less compute.

Guan Wang, Changling Liu, Chenyu Wang, Cai Zhou and 5 more

Published May 20, 2026 · 0 citations · ▲ 322 on Hugging Face · Code ★ 2,134

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models

DataFlex unifies sample selection, mixture adjustment, and reweighting for LLMs via a modular LLaMA-Factory framework that improves MMLU and perplexity with faster runtimes.

Hao Liang, Zhengyang Zhao, Mingrui Chen, Meiyi Qiang and 21 more

Published Mar 27, 2026 · 0 citations · ▲ 279 on Hugging Face · Code ★ 2,958

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Attention Residuals

Attention Residuals replace fixed residual accumulation with softmax attention over previous layer outputs for selective, input-dependent aggregation, improving scaling and downstream performance with minimal overhead.

Kimi Team, Guangyu Chen, Yu Zhang, Jianlin Su and 33 more

Published Mar 16, 2026 · 0 citations · ▲ 198 on Hugging Face · Code ★ 3,515

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

OPUS defines optimizer-induced update-space data utility for dynamic LLM pre-training selection, outperforming full-scale baselines with minimal overhead.

Shaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu and 8 more

Published Feb 5, 2026 · 0 citations · ▲ 354 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM𝛥 Integration into Upcycled MoE

The method expands multilingual LLMs via post-training PARAMΔ integration into upcycled MoE for data-efficient language acquisition.

Hao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She and 5 more

Published 2026 · 0 citations

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Beyond the Permutation Symmetry of Transformers: The Role of Rotation for Model Fusion

Rotation symmetry generalizes permutation symmetry continuously for transformers, improving parameter matching and model fusion across language and vision tasks.

Binchi Zhang, Zaiyi Zheng, Zhengzhang Chen, Jundong Li

Published Feb 1, 2025 · 0 citations

– ReadersNo votes yet. 2 from authors or colleagues not counted
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Uncovering Scaling Laws for Large Language Models via Inverse Problems

Inverse problem methods uncover scaling laws for large language models, revealing predictive relationships between model size, data, and performance from abstract evidence.

Arun Verma, Zhaoxuan Wu, Zijian Zhou, Xiaoqiang Lin and 14 more

Published 2025 · 0 citations

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Position Paper: Data-Centric AI in the Age of Large Language Models

This position paper argues for data-centric AI in the age of large language models and proposes research directions.

Xinyi Xu, Zhaoxuan Wu, Rui Qiao, Arun Verma and 15 more

Published 2024 · 2 citations

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

LMBot: Distilling Graph Knowledge into Language Model for Graph-less Deployment in Twitter Bot Detection

LMBot distills graph neural network knowledge into language models for efficient graph-less Twitter bot detection, achieving state-of-the-art results across four benchmarks.

Cai, Zijian, Zhaoxuan Tan, Zhenyu Lei, Zhu, Zifeng and 3 more

Published Jun 30, 2023 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

BERT pre-trains deep bidirectional Transformers using masked language modeling and next sentence prediction, achieving state-of-the-art results across language understanding tasks via fine-tuning.

Jacob Devlin, Ming‐Wei Chang, Kenton Lee, Kristina Toutanova

Published 2019 · 34,030 citations

– ReadersNo votes yet
16/21 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Rethinking Language Model Scaling under Transferable Hypersphere Optimization

Liliang Ren, Yang Liu, yelong shen, Weizhu Chen

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Language Models Need Sleep

Sangyun Lee, Sean McLeish, Tom Goldstein, Giulia Fanti

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Meta-memorization and memorization scaling laws in transformers

Alex Nguyen, Kenneth Norman, Gautam Reddy

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Express Language Modeling

Albert Gong, Annabelle M Carrell, Raaz Dwivedi, Lester Mackey

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Attention Is More Than You Need: Spectral Redundancy in Multi-Head Routing

Charles Cao

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Learning Rate Transfer in Normalized Transformers

Boris Shigida, Boris Hanin, Andrey Gromov

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Position Without Positional Embeddings: A Directional Mechanism in NoPE Transformers

Matan Avitan, Ido Nachum, Yanai Elazar, Yoav Goldberg

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Towards Closing the Autoregressive Gap in Language Modeling via Entropy-Gated Continuous Bitstream Diffusion

Georgios Batzolis, Mark Girolami, Luca Ambrogioni

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Closing the Loop with Fixed-Point Self-Attention

Mrinal Mathur, Barak Pearlmutter, Sergey Plis

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

On the Origin of Algorithmic Progress in AI: Evidence from Language Model Pre-Training

Hans Gundlach, Alex Fogelson, Jayson Lynch, Ana Trisovic and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Knocking-Heads Attention: Drop-in Shared Projections for Cross-Head Coordination

Zhanchao Zhou, Xiaodong Chen, Haoxing Chen, Zhenzhong Lan and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Pretraining Data Statistics Shape the Phases of Learning Entity Comparison in Language Models

Yik S Chan, Jing Huang, Yanai Elazar, Atticus Geiger

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Self-Distillation of Hidden Layers for Masked Self-Supervised Representation Learning

Scott C. Lowe, Anthony Fuller, Sageev Oore, Evan Shelhamer and 1 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

LIU Hanqing, Jianjun Cao, Yuanze Li, Zijian Zhou

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

On the Information Loss of Multi-Token Prediction: Origin and Solution

Tangyu Jiang, Zhanke Zhou, Haodi Wang, Yiu-ming Cheung and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Geometrically Disentangling Concept Learning from the Language Modeling Loss

Yupei Wang, Neil R Mallinar, Misha Belkin, Alex Warstadt

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

PipeFSDP: Efficient Pipeline Parallel under Fully Sharded Data Parallel for Large Language Model Training

Xinglin Pan, Mingji Han, Penghao Zhao, Lin Zheng and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Towards Precise Knowledge Distillation for Large Language Models via Knowledge Probing

Jiajun Liu, Yao He, Wenjun Ke, Peng Wang and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

PGSB: Pretrained-Guided Shared Basis for LoRA Model Merging

Muqing Liu, Chongjie Si, Zhuoya Liu, Yuheng Jia

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Rethinking Softmax Attention: Polynomial Activations for Transformers

Hemanth Saratchandran, Jianqiao Zheng, Yiping Ji, Wenbo Zhang and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

One Temperature to Rule Them All?

Aviv Orly, Ori Shem Ur, Yaron Oz

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Harder the Task, Sparser the Representation: Sparsity as a Learning Signature of Capability in LLMs

Mingyu Jin, Yutong Yin, Jingcheng Niu, Qingcheng Zeng and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

First-Token Attraction in Mamba Dynamics

Trinh Nguyen, Duy-Tung Pham, Minh-Khoi Nguyen-Nhat, Hoang-Son Do and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

The Cost of Absolute Position: A Spread-Expressivity Tradeoff for Additive Positional Encodings

Noah Mitchell, Isaac Gabriel, Alexander Wyatt

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Circuit-Level Knowledge Distillation for Large Language Models

Yashuo Luo, Tongxu Wang, Siyuan He, Chunyu Wei

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Virtual Head Attention

Xibo Ding, Guoxia Wang, Jinle Zeng, Jiabin Yang and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Inside Emergence: Structure-Behaviour Gaps in Language Model Training

Hak Hyun Kim, Yash Raj, Soroush Vosoughi

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Decomposing and Reshaping Scaling Laws through Token Learning Times

Pingjie Wang, zechenhu, Peiru Yang, Jingtao Han and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Contribution-Aware Structured Sparsity for Model Merging

Yan Li, GUIPING CAO, Meng XU, Tao Jiang and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Speculative Self-Distillation enables Efficient Knowledge Internalization

Shayan Talaei, Agam Bhatia, Arshia Soltani Moakhar, Jonas Hübotter and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SAGE: Semantically Disentangled Representation Learning through Latent Geometry Constraint and Large Language Model

Qiuyu Chen, Liang Xu, Yunnan Wang, Mingqi Yuan and 8 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

A Simple Unigram Cross-Entropy Lens on the Lexical Imprint of Pre-training Data

Jeonghoon Kim, Woojin Chung, Woomin Song, Cheonbok Park and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DepthGraft: Structural Regularization through Hierarchical Cross-Layer KV Reconstruction

Xiaohan Qin, Xiangdong Zhang, Yu Wang, Huaijin Wu and 3 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

ExpertNavigator: Functionally Coherent Expert Grouping and Pairwise-Ranked Routing for High-Fidelity Dense-to-MoE Conversion

Xiaohan Qin, Cancheng Zhang, Xiaoxing Wang, Xiangdong Zhang and 3 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Why Routers Freeze: Infinite Width Learning Dynamics for Mixture of Experts

Anish Dhir, Volkan Cevher, Leena Chennuru Vankadara

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

FineRMoE: Dimension Expansion for Finer-Grained Expert with an Upcycling Approach

Ning Liao, Xiaoxing Wang, Xiaohan Qin, Junchi Yan

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

\texttt{FEROM}: Frontier Endogenous Reveal-Order Marginal Policy Optimization for Masked Diffusion LMs

Zian Su, Ziyang Huang, Kaiyuan Zhang, Xiangyu Zhang

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Mitigating Memorization Where It Happens

Xuanqi Zhang, HaoYang Shang, Xiaoxiao Li

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models

Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Language Models Demand New Computer Science

Chang Yang, Xinrun Wang, Shuxin Li, Qinggang Zhang and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Better Language Models Require Better Domain-Specific Inductive Biases

Damien Teney, Liangze Jiang, Zachary Shinnick, Hemanth Saratchandran and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

From Infrastructure to Interface, the AI Value Chain Drives LLM Homogenization

Khaoula Chehbouni, Cléa Chataigner, Prakhar Ganesh, Pablo Piantanida and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

HetCCL: Efficient LLM Training on Heterogeneous Vendor GPUs

Heehoon Kim, Chris Jaehwan Lee, Taejeoung Kim, Jongwon Park and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Quantile Benchmarking Heterogeneous Web Corpora in Open LLM Pretraining

Kairong Luo, Zhenbo Sun, Xinyu Shi, Shengqi Chen and 8 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Beyond the Node Barrier: Zero-Shot Strategy Planning for LLM Training on Super-Nodes

Shijie Shen, Chong Li, Pierre Leca, Jiong Lou and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

When is Warmstarting Effective for Scaling Language Models?

Neeratyoy Mallik, Maciej Janowski, Johannes Hog, Herilalaina Rakotoarison and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition

Yuxuan Lou, Kai Yang, Geng Zhang, Yong Liu and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Rethinking Bayesian Optimization for Co-Optimizing LLM Training Configurations

Zhiliang Chen, Alfred Leong, Shao Yong Ong, Apivich Hemachandra and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Witness Overlap: Directional Provenance Inside Open-Weight Model Families

Siyuan Li, Haoxuan Zeng, Xin Luo, Fernando Jia and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Language Denoising Objectives Extend the Value of Limited Data

Justin Lovelace, Christian Belardi, Srivatsa R Kundurthy, Shriya Sudhakar and 1 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

ClusQuant: Mitigating Outliers with Clustering-Based Representations for Low-Precision LRMs

Xingyu Liu, Xiangyang Yin, Tianhua Xia, Haiyu Wang and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Looped Diffusion Language Models

Selective layer looping improves masked diffusion model training efficiency and reasoning performance via depth scaling without added parameters and flexible inference compute scaling.

Sanghyun Lee, Chunsan Hong, Seungryong Kim, Jonghyun Lee and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Why Muon Outperforms Adam: A Curvature Perspective

Muon achieves larger one-step loss decreases than Adam via lower curvature penalties driven by reduced normalized directional sharpness rather than update scale, with advantages amplified by data imbalance and within-layer curvature.

Shuche Wang, Fengzhuo Zhang, Jiaxiang Li, Dirk Bergemann and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 10 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs

HESTIA replaces hard quantization with Hessian-guided temperature-controlled soft relaxation for low-bit LLM training, improving 1.58-bit Llama-3.2 1B and 3B zero-shot accuracy by 5.39% and 4.34%.

Guoan Wang, Feiyu Wang, Zongwei Lv, Yikun Zong and 2 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models

PACE wraps AdamW to pull live weights toward their EMA with clipped per-coordinate control, improving iterate-averaged LM training and reducing error by arbitrarily large factors in quadratics while boosting 1-2B SFT and GPT-2 pretraining.

Kwok C Au, Adam Block

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Discrete Stochastic Localization for Non-autoregressive Generation

Discrete stochastic localization uses unit-sphere embeddings with SNR-invariant denoisers to enable flexible continuous-state non-autoregressive generation with improved fidelity and hybrid sampling.

Yunshu Wu, Jiayi Cheng, Longxuan YU, Partha Thakuria and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

A Theory of Training Profit-Optimal LLMs

Economic model combining scaling laws with microeconomics shows profit-optimal LLM training scales near-linearly with hardware efficiency and sub-quadratically in cost, while current expenditure trends are only optimal under compute-bound assumptions.

Sophie Hao, Will Merrill

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

LittleLearner: Language Models Under Pedagogically-Controlled Knowledge Exposure

LittleLearner is a 5B-parameter model trained on grade-capped elementary data to study controlled knowledge acquisition and bounded capability growth.

Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws

Masked-input regularization improves autoregressive pretraining over weight decay alone, and SoftQ scaling laws better capture data-constrained training than Chinchilla.

Zhiwei Xu, Shihao Wu, Hanseul Cho, Wei Hu and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs

Ghosted Layers recovers layer-pruned LLMs via training-free activation alignment with a closed-form linear operator, outperforming constrained baselines.

Vincent-Daniel (Juyoung) Yun, Junhyuk Jo, Sai Praneeth Karimireddy, Sunwoo Lee

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

xHC: Expanded Hyper-Connections

xHC expands Transformer hyper-connections beyond four streams via sparse updates and temporal augmentation, improving scaling efficiency. It boosts 18B MoE downstream scores by 4.0 points over mHC with lower compute and reduced memory traffic via xHC-Flash.

Xiangdong Zhang, Xiaohan Qin, Tuo Dai, Xiaoming Shi and 7 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 54 on Hugging Face · Code ★ 68

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 3/5
83%Must read
?Must readVote to see the score

dMoE: dLLMs with Learnable Block Experts

dMoE aggregates token-level expert distributions into block-level routing for diffusion LLMs, reducing uniquely activated experts from 69.5 to 14.6 while retaining 99.11% performance and cutting memory use up to 79.84%.

Sicheng Feng, Zigeng Chen, Gongfan Fang, Xinyin Ma and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 32 on Hugging Face · Code ★ 48

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Flow Map Language Models: One-step Language Modeling via Continuous Denoising

Continuous flow language models outperform discrete diffusion in quality and speed, and distilling their unique flow map enables one-step generation surpassing eight-step discrete diffusion.

Chanhyuk Lee, Jaehoon Yoo, Manan Agarwal, Sheel Shah and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 4 on Hugging Face · Code ★ 172

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

RePlaid, a continuous diffusion language model aligned with modern discrete architectures, achieves scaling laws rivaling discrete diffusion and sets a continuous diffusion perplexity record of 22.1 on OpenWebText.

Zhihan Yang, Wei Guo, Shuibai Zhang, Subham Sahoo and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

DiLaDiff: Distilled Latent-augmented Diffusion for Language Modeling

DiLaDiff proposes a latent-augmented masked diffusion language model with consistency distillation that improves quality and accelerates inference by generating continuous latents in negligible time.

Jean-Marie Lemercier, Tomas Geffner, Morteza Mardani, Karsten Kreis and 2 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Leviathan: Decoupling Input and Output Representations in Language Models

Leviathan decouples input and output embeddings via learned continuous token vectorization, cutting perplexity up to 9% and rare-token perplexity by 81% with only 0.2% extra parameters.

Reza T Batley, Sourav Saha

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory

Memory Grafting uses frozen hidden states from a grafting model as conditional n-gram memory for language models, improving average benchmarks to 53.86 versus 52.43 for vanilla Engram at 2.8B scale with minimal overhead.

Runxi Cheng, Yuchen Guan, Yongxian Wei, Qianpu Sun and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation

Pion updates weight matrices via orthogonal equivalence transformations that preserve singular values during LLM training, offering a stable, competitive optimizer alternative.

Kexuan Shi, Hanxuan Li, Zeju Qiu, Yandong Wen and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face · Code ★ 39

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Accelerating LLM Pre-Training through Flat-Direction Dynamics Enhancement

A Riemannian ODE framework shows adaptive optimizers dampen flat directions too conservatively, and LITE accelerates Muon and SOAP by boosting flat-direction updates, cutting LLM pre-training time.

Shuchen Zhu, Rizhen Hu, Mingze Wang, Mou Sun and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts

Geometric coupling aligns router and expert gradients along shared input directions in sparse Mixture-of-Experts, while load-balancing losses disrupt it, and cosine-similarity routing achieves low imbalance with minimal perplexity cost.

Sagi Ahrac, Noya Hochwald, Mor Geva

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Sparse Layers are Critical to Scaling Looped Language Models

Looped-MoE models scale better than standard transformers via routing divergence that recovers expressivity, and loop boundaries enable efficient early exits with minimal quality loss.

Ryan Lee, Jacob Biloki, Edward J Hu, Jonathan May

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

From Collapse to Improvement: Statistical Perspectives on the Evolutionary Dynamics of Iterative Training on Contaminated Sources

Statistical analysis shows iterative training on contaminated synthetic data avoids model collapse and recovers the true distribution with sufficient fresh samples and appropriate mixture weights.

Soham Bakshi, Sunrit Chakraborty

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

PithTrain: A Compact and Agent-Native MoE Training System

PithTrain is a compact agent-native MoE training framework that matches production throughput while reducing agent turns by 62% and GPU time by 64% on framework tasks.

Ruihang Lai, Hao Kang, Haozhan Tang, Akaash R Parthasarathy and 5 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

TIDE: Every Layer Knows the Token Beneath the Context

TIDE injects token identity into every layer via EmbeddingMemory to fix rare-token undertraining and contextual collapse, improving language modeling and downstream performance.

AJAY JAISWAL, Lauren Hannah, Han-Byul Kim, Duc Hoang and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Bug or Feature$^2$: Weight Drift, Activation Sparsity, and Spikes

Standard losses and biased activations induce negative weight drift that drives early training dynamics and extreme sparsity across architectures, with squared activations sharply improving accuracy until a cliff near 70% sparsity unless clipping controls intermediate spikes.

Egor Shvetsov, Aleksandr Serkov, Shokorov Viacheslav, Redko Dmitry and 2 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026 · ▲ 1 on Hugging Face · Code ★ 1

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Towards Understanding Self-Pretraining for Sequence Classification

Self-pretraining improves sequence classification by learning proximity-biased attention patterns that supervised training misses from random initialization.

Omar Coser, Loredana Zollo, Paolo Soda, Antonio Orvieto

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Support Before Frequency in Discrete Diffusion

Discrete diffusion models learn data support before frequencies because reverse edits scale by validity first and coefficients second; absorbing diffusion prioritizes validity-improving moves over uniform diffusion's trichotomy.

Adrian Müller, Antoine Gonon, Zebang Shen, Ya-Ping Hsieh and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Swimba: Switch Mamba Model Scales State Space Models

Swimba routes expert SSM streams via parameter-space MoE to scale selective state space model capacity without increasing recurrent state update costs. Under matched FLOPs, it achieves slightly better average performance with minor latency and throughput trade-offs.

Zhixu Du, Krishna Teja Chitty-Venkata, Murali Emani, Venkatram Vishwanath and 2 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Bridging Compute- and Data-Optimal Pretraining

A unified compute-data scaling framework introduces token effectiveness to bridge data-rich and fixed-corpus regimes, showing diminishing returns and three operational regimes that render classical compute-optimal allocations suboptimal.

Tian Qin, Kimia Hamidieh, David Alvarez-Melis

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
89%Must read
?Must readVote to see the score

Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization

A framework maps compute budgets to optimal Mixture-of-Experts architectures via joint FLOP, active, and total parameter constraints, yielding robust scaling laws across hundreds of models with widening near-optimal flexibility at scale.

Weilin Wan, Jingtao Han, Debing Zhang, Weizhong Zhang and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5
89%Must read
?Must readVote to see the score

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation

RISE sketches LLM output-layer influence hotspots into compressed dual-channel sketches, reducing storage up to 112x versus gradient methods while scaling to 32B parameters for attribution and data valuation.

yide ran, Jianwen Xie, Minghui Wang, W. Jim Zheng and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
89%Must read
?Must readVote to see the score

Training Language Models via Neural Cellular Automata

Neural cellular automata generate synthetic pre-training data that improves language model convergence and downstream reasoning faster than natural text.

Dan Lee, Seungwook Han, Akarsh Kumar, Pulkit Agrawal

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 8 on Hugging Face · Code ★ 88

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
70%Highly rated
?Highly ratedVote to see the score

Progressive Residual Warmup for Language Model Pretraining

ProRes progressively warms up deeper layer residuals to stabilize pretraining, accelerating convergence and improving generalization across model scales.

Tianhao Chen, Xin Xu, Lu Yin, Hao CHEN and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 36 on Hugging Face · Code ★ 8

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 5 of 20 reviewers recommend it
lenient 3/5
medium 2/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup

MoME replaces fixed token memory entries with context-gated mixtures of slots to disambiguate polysemous tokens, improving sparse lookup performance across small-scale language model backbones with favorable scaling.

Muchen Li, Leonid Sigal, Renjie Liao

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 12 on Hugging Face · Code ★ 7

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
89%Must read
?Must readVote to see the score

Towards Scalable Data Diversification for Language Model Pretraining via Leverage Score Sampling

Leverage Score Sampling enables scalable data diversification for LM pretraining via leverage scores, improving diversity by 9.2% and speed by 72×.

Zailin Ma, Quzhe Huang, Yujun Li, Congyuan Rao and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Benchmarking Optimizers for Large Language Model Pretraining

Standardized LLM pretraining benchmarks compare optimizers across model sizes, batch sizes, and training durations to guide selection and highlight future research directions.

Andrei Semenov, Matteo Pagliardini, Martin Jaggi

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 25 on Hugging Face · Code ★ 61

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
83%Must read
?Must readVote to see the score

It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

SYNTH is an open-source synthetic dataset derived from Wikipedia that collapses pre-, mid-, and post-training into one stage, training competitive small models with 10-140x fewer tokens and higher factual precision than web-crawled data.

Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov and 6 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
83%Must read
?Must readVote to see the score

TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation

TransitLM provides 13 million transit records to train LLMs for map-free route generation, yielding accurate routes that implicitly ground GPS coordinates without maps.

Hanyu Guo, JieDong Yang, Chao Chen, Longfei Xu and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 127 on Hugging Face · Code ★ 129

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence

Synthetic noise in pretraining data causes LLM loss divergence with probability scaling by noise type, amount, and model size, exhibiting activation patterns distinct from high-learning-rate failures.

Qizhen (Irene) Zhang, Ankush Garg, Jakob Foerster, Niladri S. Chatterji and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 2/5
80%Must read
?Must readVote to see the score

Communication-Efficient LLM Adaptation over Decentralized GPU Meshes

An asynchronous two-circuit system with spectral correction enables compressed LLM adaptation over decentralized GPUs, yielding up to 40× speedups with dense-level accuracy.

Sameera Ramasinghe, Shamane Siriwardhana, Thalaiyasingam Ajanthan, Hadi Mohaghegh Dolatabadi and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Continuous Latent Diffusion Language Model

Cola DLM is a hierarchical latent diffusion language model using continuous latent priors and block-causal DiTs to achieve scalable non-autoregressive text generation. It demonstrates strong scaling behavior and quality competitive with matched autoregressive baselines across benchmarks.

Hongcan Guo, Qinyu Zhao, Yian Zhao, Shen Nie and 7 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 86 on Hugging Face · Code ★ 293

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 1/5
86%Must read
?Must readVote to see the score

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization

Researchers derive maximally scale-stable parameterizations for Mixture-of-Experts via dynamical mean-field theory, yielding robust learning-rate transfer and monotonic scaling gains across regimes.

Leena Chennuru Vankadara, Moritz Haas, Luke Hayward, Sebastian Bordt and 1 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026 · ▲ 1 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 3/5
80%Must read
?Must readVote to see the score

CausalMix: Data Mixture as Causal Inference for Language Model Training

CausalMix frames data mixture optimization as causal inference to dynamically estimate optimal mixtures via conditional average treatment effects, improving LLM performance without retraining proxy models.

Zinan Tang, YUKUN ZHANG, Shaomian Zheng, Zhuoshi Pan and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 19 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
88%Must read
?Must readVote to see the score

LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling

LangFlow closes the continuous-discrete language-modeling gap via flow matching and a learnable noise schedule, matching discrete diffusion perplexity and exceeding autoregressive zero-shot results on four benchmarks.

Yuxin Chen, Chumeng Liang, Hangke Sui, Ruihan Guo and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 15 on Hugging Face · Code ★ 96

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 3/5
medium 10/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

A Bitter Lesson for Data Filtering

In high-compute, data-scarce pretraining, removing filters and using all data improves large models by letting them learn from nominally poor data.

Christopher Mohri, John Duchi, Tatsunori Hashimoto

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
80%Must read
?Must readVote to see the score

Neural Neural Scaling Laws

NeuNeu predicts downstream scaling via time-series extrapolation of task accuracies and token-level losses, achieving 1.99% MAE and 44% lower error than logistic scaling laws with zero-shot generalization.

Michael Hu, Jane Pan, Ayush Rajesh Jhaveri, Nicholas Lourie and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

Synthetic pre-pre-training improves language model robustness to noisy pre-training data by inhibiting noise self-modeling and reducing required natural-text tokens by up to 49%.

Xu Guo, Runyu Peng, Jian Tong, Yunhua Zhou and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Scaling Linear Mode Connectivity and Merging to Billion Parameter Pretrained Transformers

A scalable dual-learning framework applies parameter symmetry alignments to enable near-barrier-free linear merging of billion-parameter pretrained transformers.

Tianyi Li, Zhiqiang Shen

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 2/5
medium 7/10
strict 2/5
83%Must read
?Must readVote to see the score

Normalized Architectures are Natively 4-Bit

Normalizing weights and hidden states to the unit hypersphere makes nGPT robust to 4-bit arithmetic, enabling stable end-to-end NVFP4 training without extra scaling or transforms.

Maxim Fishman, Brian Chmiel, Ron Banner, Daniel Soudry and 1 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
86%Must read
?Must readVote to see the score

Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation

Trajectory-Shaped Discrete Flow Matching guides discrete flow matching training via an energy-based midpoint evaluator, letting small students outperform large teachers with 32% lower perplexity at 128x speed.

Amin Karimi Monsefi, Dominic Culver, Nikhil Bhendawade, Manuel R Ciosici and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 3 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
89%Must read
?Must readVote to see the score

Next-Latent Prediction Transformers Learn Compact World Models

NextLat adds latent self-prediction to transformers, theoretically converging to belief states and empirically improving world modeling, reasoning, and inference speed.

Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward Hu and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face · Code ★ 196

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 2/5
83%Must read
?Must readVote to see the score

Truth as a Compression Artifact in Language Model Training

Language models prefer correct answers because errors are less compressible, not because they detect truth directly, with coherent false rules eliminating accuracy until competing rules restore it.

Konstantin Krestnikov

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 3/5