Good Papers

Geometric Factual Recall in Transformers

Transformers memorize facts geometrically via linear superpositions and MLP selectors, needing only logarithmic dimensions and enabling zero-shot MLP transfer.

Shauli Ravfogel, Gilad Yehudai, Joan Bruna, Alberto Bietti

Published 2026Paris Poster Session 3 · Thu, Dec 10, 12:30 PM–2:30 PM local time · Paris Poster Hall▲ 1 on Hugging FacearXiv ↗OpenReview ↗

88%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel15/20reviewers recommend it
lenient 4/5
medium 7/10
strict 4/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
A striking geometric account proves logarithmic embeddings and relation-conditioned MLP selectors can memorize facts with zero-shot transfer, though its synthetic single-layer bijections leave real-world multi-hop validation and production latency uncertain.

Abstract

How do transformer language models memorize factual associations? A common view casts internal weight matrices as associative memories over pairs of embeddings, requiring parameter counts that scale linearly with the number of facts. We develop a theoretical and empirical account of an alternative, \emph{geometric} form of memorization in which learned embeddings encode relational structure directly, and the MLP plays a qualitatively different role. In a controlled setting where a single-layer transformer must memorize random bijections from subjects to a shared attribute set, we prove that a logarithmic embedding dimension suffices: subject embeddings encode \emph{linear superpositions} of their associated attribute vectors, and a small MLP acts as a relation-conditioned selector that extracts the relevant attribute via ReLU gating, and not as an associative key-value mapping. We extend these results to the multi-hop setting -- chains of relational queries such as ``Who is the mother of the wife of $x$?'' -- providing constructions with and without chain-of-thought that exhibit a provable capacity-depth tradeoff, complemented by a matching information-theoretic lower bound. Empirically, gradient descent discovers solutions with precisely the predicted structure. Once trained, the MLP transfers zero-shot to entirely new bijections when subject embeddings are appropriately re-initialized, revealing that it has learned a generic selection mechanism rather than memorized any particular set of facts.