Good Papers

Calibration-risk routing for controlled world-model adaptation

MC-WM partitions target data to select lower-calibration-risk world models and weights imagined policy updates via learned confidence, evaluated across 541 MuJoCo shift executions.

Yifan F. Zhang, Liang Zheng

Published Oct 1, 2026arXiv ↗

71%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel7/20reviewers recommend it
lenient 3/5
medium 3/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Calibration-risk routing cleanly partitions fit, selection, and calibration to choose lower-risk world-model families with elegant validity predicates, though 541 runs may obscure a brittle routing rule hidden by one repeated artifact-gate cell.

Abstract

Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which separates initial target data into disjoint fit, selection, and calibration partitions and deploys the family with lower standardized calibration risk. A learned confidence signal and deterministic validity predicates weight one-step imagined policy updates without rewriting physical rewards. We evaluate 540 unique reported run cells across three controlled Multi-Joint dynamics with Contact (MuJoCo) shifts; one exact-routing cell was repeated after a pre-deployment artifact gate, giving 541 completed executions.