Good Papers

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

BDH-CQ combines in-context learning with recurrent latent reasoning, achieving 29.5% ARC-AGI-1 pass@2 at $0.0007 per task to set a new cost-efficiency frontier.

Björn Engdahl, Adrian Kosowski, Jan Chorowski, Zuzanna Stamirowska, Przemysław Uznański, Junlin Jiang, Rohan Phadke, Remigiusz Kinas, Richard Zhong

Published Aug 10, 2026▲ 797 on Hugging FaceCode ★ 11,072arXiv ↗

71%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel6/20reviewers recommend it
lenient 3/5
medium 3/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
BDH-CQ establishes a striking cost-accuracy frontier with 29.5% pass@2 at $0.0007/task via recurrent latent reasoning, yet its single-point evaluation lacks variance reporting, ablations isolating recurrence from scale, and rigorous validation of its nonverbal inference mechanism.

Abstract

We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model's recurrent memory; the model then solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning. We evaluate the model on the public ARC-AGI-1 evaluation set and use controlled ARC-like interventions to study what it learns from demonstrations, how consistently it applies an inferred transformation, and which concepts remain difficult. A 150M-parameter configuration reaches 29.5% pass@2 at a computed inference cost of \$0.0007 per task. This operating point breaks through the previously reported ARC-AGI-1 cost-accuracy Pareto frontier, establishing a new state of the art in benchmark cost efficiency.