Good Papers

Showing Training dynamics Show all papers

45%Niche pick
?Niche pickVote to see the score

Two-Clustering Regime of Token Dynamics in Causal Attention

Trinh Nguyen, Duy-Tung Pham, Hoang-Son Do, Tan Nguyen and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Information Propagation via Sign-Flip Dynamics

Hyunwoo Lee, Hyojae Lim, Dohyun Kwon

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent

A representation-readout decomposition attributes grokking and double descent to competing encoder and classifier dynamics, showing delayed generalization stems from gradual representation learning rather than lazy-to-rich transitions.

Chi-Ning Chou, Oscar Uzdelewicz, Neng-Chun Chiu, Yao-Yuan Yang and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

Why Geometric Continuity Emerges in Deep Neural Networks: Residual Connections and Rotational Symmetry Breaking

Residual connections and symmetry-breaking nonlinearities cause geometric continuity across deep network layers, with activation and normalization distributing it differently across singular directions and projection types.

Kyungwon Jeong, Won-Gi Paeng, Honggyo Suh

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

The Geometric Inductive Bias of Grokking: Bypassing Phase Transitions via Architectural Topology

Architectural topology modifications eliminate Transformer's grokking phase by bounding representations and fixing attention, but only when aligned with task symmetries.

Alper YILDIRIM

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 2/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models

Non-commuting SGD updates leave localized, steerable parametric memory of training order captured by gradient Lie brackets and readable via sparse vocabulary projections that identify model origins with 92% accuracy.

John Sweeney

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 2/5
medium 6/10
strict 4/5
76%Highly rated
?Highly ratedVote to see the score

From Density Matrices to Phase Transitions in Deep Learning: Spectral Early Warnings and Interpretability

A 2-datapoint reduced density matrix provides unified spectral early warnings of training phase transitions and interpretable eigenvectors across deep learning settings.

Max Hennick, Guillaume Corlouer

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Estimating Implicit Regularization in Deep Learning

Gradient matching methods empirically estimate implicit regularization in deep networks, recovering explicit penalties and revealing dropout's implicit L2 effects.

Joseph H Rudoler, Kevin Tan, Giles Hooker, Konrad Kording

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5