Good Papers

Rethinking Expressivity and Efficiency in Test-Time Training

E²-TTT derives a closed-form chunk-level state transition that exactly reproduces per-token update dynamics, enabling parallel training that retains temporal structure, matches chunk-wise throughput, and achieves over 90% needle-in-a-haystack accuracy at 8× training length.

Zeyun Zhong, Joya Chen, Manuel Martin, Frederik Diederichs, Jürgen Gall, Jürgen Beyerer

Published 2026Paris Poster Session 4 · Thu, Dec 10, 5:30 PM–7:30 PM local time · Paris Poster Hall▲ 2 on Hugging FaceCode ★ 4arXiv ↗OpenReview ↗

83%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel13/20reviewers recommend it
lenient 3/5
medium 9/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Test-Time Training (TTT) enables long-context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per-token update dynamics with the hardware efficiency of chunk-wise approximations. We propose E$^2$-TTT (Expressive and Efficient TTT) to bridge this gap. Under the standard approximation of taking gradients at the chunk-start weights, we derive a closed-form state transition that exactly reproduces the chunk-end fast-weight and momentum states of the per-token recurrence. This enables fully parallelized chunk-level training while preserving the temporal structure of the update rule that prior chunk-wise methods discard. We validate E$^2$-TTT by training models up to 1.3B parameters from scratch. It performs on par with previous TTT and hybrid attention baselines in language modeling while outperforming them on in-context retrieval. Its advantage is most pronounced in length extrapolation: on the standard ``Needle in a Haystack'' passkey test, it retains over 90% accuracy at $8\times$ the training context length. Meanwhile, E$^2$-TTT can match the training throughput of efficient chunk-wise methods, demonstrating that it effectively reconciles expressivity with efficiency. The code is available at https://github.com/zeyun-zhong/E2-TTT.