This paper models parallel inference-time reasoning via particle filtering, deriving non-asymptotic guarantees and fundamental limits for sequential Monte Carlo with process reward models.
The paper generalizes neural-transport-accelerated free energy estimation to arbitrary state spaces, validating it across discrete, multimodal, and autoregressive settings while establishing group-theoretic identities linking time reversal and Doob's transforms.