
Depth-recurrent attention mixtures (Dreamer) combine sequence, depth, and sparse expert attention to scale latent reasoning efficiently, requiring 2, 8x fewer training tokens than matched baselines while improving expert diversity.
Jonas Knupp, Jan Metzen, Jeremias Bohn, Georg Groh and 1 more
Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face
– ReadersNo votes yet
11/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 11 of 20 reviewers recommend it
lenient 2/5
medium 8/10
strict 1/5