Good Papers

What Matters for Latent Reasoning with Flow Matching

FLaRe uses flow matching for latent reasoning that is useful, diverse, explainable, refinable and efficient, reaching 97% of explicit chain-of-thought accuracy at 25% latency.

Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

Published Oct 5, 2026▲ 10 on Hugging FacearXiv ↗

83%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel13/20reviewers recommend it
lenient 4/5
medium 9/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
FLaRe delivers impressive speed and accuracy gains through flow-matched latent reasoning with verified-thought refinement, yet critics remain unconvinced the flow learns genuine reasoning rather than interpolating distilled explicit chains, and the hidden verification and training costs…

Abstract

Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows, refinable with more inference compute, and efficient, costing less than an explicit CoT at comparable accuracy. Current methods rarely meet these requirements: they learn shortcuts from the question, distill the explicit CoT into their weights, or imitate it one token at a time. We focus on flow matching in a learned latent space, the family we argue is best placed to meet them, and identify the training choices that make it work. The result is Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts. A probe for each requirement shows that FLaRe improves on prior latent methods in all five. It also compares favorably with them on arithmetic benchmarks, while reaching 97% of the accuracy of explicit CoT at a quarter of its latency.