Good Papers

Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

FailBank turns runtime shield feedback into persistent VLA policy updates via failure-bank self-evolution, raising success rates up to 25.4 points and cutting policy-induced cost up to 35.6%.

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang

Published Sep 30, 2026▲ 15 on Hugging FaceCode ★ 2arXiv ↗

88%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel15/20reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
FailBank turns CBF runtime corrections into durable LoRA policy gains with counterfactual targets and quiet anchors, but its single-arena, two-backbone evaluation and unverified open-source code leave real-time deployment and cross-robot transfer unresolved.

Abstract

Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent policy-shield mismatch that blocks task progress. To address this challenge, we introduce FailBank, a four-stage self-evolving framework that converts runtime feedback into persistent policy improvement. During collection, a fixed CBF-based safety module serves as an observe-only teacher, producing counterfactual corrections while the policy remains in control. Outcome-aware admission then converts useful proposals into corrective targets and retains successful uncorrected actions as quiet anchors for guarded LoRA updates. We evaluate FailBank on the VLA-Arena benchmark across two difficulty levels and two VLA backbones. Compared with the base policies, FailBank improves the joint success-cost operating point. Across the two backbones, FailBank improves task success rate by 8.5 and 6.9 percentage points, while reducing policy-induced cumulative cost by 35.6\% and 23.8\%, respectively. Compared with runtime shielding, FailBank raises task success rate by 25.4 and 9.5 percentage points, while maintaining comparable policy-induced cumulative cost. These results show that runtime feedback can serve as persistent policy supervision rather than only as a temporary action constraint.