Good Papers

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

VeriFine co-evolves policies, curricula, and judges via adaptive diagnosis and coactive calibration to sustain embodied reasoning self-improvement. Experiments on driving and navigation show continuous gains in both policy and judge capability.

Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao, Chensheng Peng, Boyi Li, Thomas Tian, Zheng Lian, Yan Wang, Jiaqi Ma, Boris Ivanovic, Marco Pavone, Wenhao Ding

Published Oct 6, 2026▲ 4 on Hugging FacearXiv ↗

78%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
?1 reader voted. Vote to see how they split.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
VeriFine offers an elegant co-evolution of policy, curriculum, and rubric judge that sustains embodied self-improvement, though it lacks hard ablation isolating the judge loop and obscures human annotation overhead behind selective-query jargon.

Abstract

Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.