Good Papers

Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics

Attention-only transformers under token corruption implement two-stage in-context empirical Bayes via depth-refined particle dynamics and skip-connection queries, enabling depth-dependent denoising without explicit noise schedules and posterior-mean convergence to Bayes-optimal predictors.

Matthew Smart, Soumya Ganguly, Nilava Metya, Alexandre V Morozov, Anirvan Sengupta

Published 2026Atlanta Poster Session 5 · Fri, Dec 11, 10:00 AM–1:00 PM local time · Hall C1arXiv ↗OpenReview ↗

71%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel6/20reviewers recommend it
lenient 2/5
medium 3/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

We study minimal attention-only transformers under all-token corruption and show they admit a two-stage empirical Bayes interpretation. A single attention step computes a kernel-weighted posterior mean with respect to the empirical distribution defined by the context. Depth refines this distribution through particle dynamics (Stage 1), while a long-range skip-connection carries the noisy input as a query for posterior inference (Stage 2), revealing distinct statistical roles for depth and attention residuals. The framework isolates a minimal setting in which the context itself induces a depth-dependent energy landscape governing in-context inference. We show that effective denoising can emerge without an explicit noise schedule: a fixed kernel bandwidth and finite integration horizon suffice, yielding a principled depth-noise relationship. We further establish a posterior-mean recovery guarantee for a class of well-behaved priors, where the empirical estimator converges to the Bayes-optimal predictor under asymptotic conditions. Connecting these dynamics to reverse-diffusion limits, our results provide a statistical interpretation of attention as in-context inference via sample-based posterior estimation, without explicit density modeling.