Good Papers

Showing Model-based RL Show all papers

71%Highly rated
?Highly ratedVote to see the score

Calibration-risk routing for controlled world-model adaptation

MC-WM partitions target data to select lower-calibration-risk world models and weights imagined policy updates via learned confidence, evaluated across 541 MuJoCo shift executions.

Yifan F. Zhang, Liang Zheng

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

Revisiting Value Iteration: Unified Analysis of Discounted and Average-Reward Cases

Arsenii Mustafin, Xinyi Sheng, Dominik Baumann

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AdaWM: Few-Shot Adaptation of World Models to Unseen Dynamical Regimes

Zian Guan, Guozheng Li, Zilun Zhang, Zecong Tang

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Objective-Aligned Amortized Inference for Offline Bayes-Adaptive MDP Model Learning

Toru Hishinuma, Kei Senda

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Amortizing Generative Guidance for Model-Based Reinforcement Learning

Xiangteng Zhang, Guojian Zhan, Likun Wang, Jingliang Duan and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

When Actions Matter: Causal Affordances for Long-Horizon Credit Assignment in World Models

Pradeep Kumar Banerjee, Frank Röder, Nihat Ay

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Generalizing Action-Conditioned Latent World Models with Video Model Rewards

Haichao Zhang, Yijiang Li, Shwai He, Van T Le and 2 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Control Under the Wrong Model Is Better Than Under the Correct One

Francesco Damiani, Ruben Moreno Bote

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

BlockFormer: Transformer-based inference from interaction maps

Eloïse Touron, Pedro Rodrigues, Julyan Arbel, Nelle Varoquaux and 1 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Bridging Risk Approximation Gaps in Model Predictive Task Sampling via In-Context Modeling

Jiarong Wen, Qi Tao, Zhang Kaiyu, Yun Qu and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

EpicWorldModel: Exploration-driven Planning with Latent World Models

Bowen Feng, Julian Ost, Zhiting Mei, Anirudha Majumdar and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Towards Optimism-Pessimism Trade-off in Model-based Offline-to-Online Reinforcement Learning

Guochen Zhou, Yijun Yang, Qiqi Duan, Qing Su and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Near-Optimal Sample Complexity of Robust Reinforcement Learning with KL Uncertainty Set

Yudan Wang, Zilong Deng, Nathaniel D Bastian, Shaofeng Zou

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Hierarchical World Models with Implicit Dynamics

Gaoyue Zhou, Yvonne Wu, Zichen Cui, Nicolas Ballas and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

K-PWM: Control-Oriented Structured World Models under Partial Observation

Santosh M Rajkumar, Sriram Narayanan, Samuel E Otto, Debdipta Goswami

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Model-Based Online Decision Making via Generative Trajectory Planning

Haldun Balim, Yilun Du, Na Li

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Joint Learning of Hierarchical Neural Options and Abstract World Model

AgentOWL jointly learns hierarchical neural options and an abstract world model for sample-efficient skill acquisition, outperforming baselines on object-centric Atari games with fewer samples and stronger generalization.

Top Piriyakulkij, Wolfgang Lehrach, Kevin Ellis, Kevin Murphy

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
70%Highly rated
?Highly ratedVote to see the score

Reinforcement Learning with Multi-Step Lookahead Information Via Adaptive Batching

Adaptive batching policies process multi-step lookahead via state-dependent batches, yielding near-optimal regret bounds for tabular reinforcement learning.

Nadav Merlis

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 4 of 20 reviewers recommend it
lenient 2/5
medium 2/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Dyna-Style Safety Augmented Reinforcement Learning: Staying Safe in the Face of Uncertainty

Dyna-SAuR learns scalable safety filters and policies via uncertainty-aware dynamics to reduce training failures by two orders of magnitude versus baselines.

Artur Eisele, Bernd Frauenknecht, Friedrich Solowjow, Sebastian Trimpe

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Planning as Dynamics Relaxation: Hippocampal Recurrent Network Realizes Optimal Goal-Directed Navigation

Hippocampal recurrent networks achieve optimal navigation via relaxation dynamics equivalent to linearly-solvable Markov decision processes.

宇航 他, Junfeng Zuo, Tianhao Chu, Si Wu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
83%Must read
?Must readVote to see the score

StarWM: Self-Supervised Trained Attention Routing for Robust World Models

StarWM uses self-supervised dynamics to route attention for selective reconstruction, yielding robust world models that preserve task-relevant states and discard distractors across video backgrounds.

Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte, Daniel Schmid and 3 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
86%Must read
?Must readVote to see the score

Inverting the Bellman Equation: From $Q$-Values to World Models

Value-based agents trained on diverse reward functions implicitly encode world models, extractable via P-learning, with sufficient conditions for exact dynamics recovery and cross-goal generalization.

Alistair Letcher, Mattie Fellows, Alexander D. Goldie, Jonathan Richens and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 3/5
74%Highly rated
?Highly ratedVote to see the score

Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees

Applying CVaR to the immediate belief cost targets per-step state uncertainty while preserving standard MDP structure, enabling any expectation-based planner to become risk-sensitive with unchanged algorithms and end-to-end finite-time guarantees.

Yaacov Pariente, Vadim Indelman

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 2/5
medium 6/10
strict 1/5
83%Must read
?Must readVote to see the score

Graph Sparse Sampling: Breaking the Curse of the Horizon in Continuous MDP Planning

Graph Sparse Sampling shares sampled futures across actions to avoid exponential horizon dependence in continuous MDP planning, with polynomial sample bounds and strong long-horizon control performance.

Idan Lev-Yehudi, Vadim Indelman

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
80%Must read
?Must readVote to see the score

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving

CoPhy distills vision-language cognition into a BEV encoder and pairs it with an auto-regressive world model for action-conditioned forecasting to enable reinforcement learning with dual physical and cognitive rewards, achieving state-of-the-art autonomous driving results.

Yang Wu, Qiang Meng, Zhaojiang Liu, Youquan Liu and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

Default Feature Representations of the Cognitive Map

Default Feature Representations parameterize predictive cognitive maps via a fixed feature basis and an adaptive environment operator, enabling rapid sample-based replanning and local grid-cell remapping.

Armin Bazarjani, Payam Piray

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret

Strong duality holds for weakly communicating average-reward CMDPs, yielding a primal-dual algorithm with O(T^{2/3}) regret and constraint violations.

Kihyun Yu, Beomhan Baek, Dabeen Lee

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 2/5
medium 5/10
strict 3/5
71%Highly rated
?Highly ratedVote to see the score

Graph-Based Stochastic-Power-UCT: Monte-Carlo Graph Search with Power Mean Estimation

GS-Power-UCT shares same-depth states in stochastic planning graphs to reuse samples while maintaining O(n^{-1/2}) convergence, with adaptive-horizon variants achieving optimal infinite-horizon values.

Tung Tran, Viet Bao, Hoang Ta, Tuan Dam

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 1/5
medium 4/10
strict 2/5
86%Must read
?Must readVote to see the score

SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning

SLOPE constructs optimistic potential landscapes via distributional regression to amplify sparse success signals and guide planning, outperforming baselines across sparse reward benchmarks.

Yao-Hui Li, Zeyu Wang, Xin Li, Wei Pang and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Predictive but Not Plannable: RC-aux for Latent World Models

RC-aux improves latent world model planning by adding multi-horizon prediction and budget-conditioned reachability supervision to align latent spaces with long-horizon search.

Wenyuan Li, Guang Li, Keisuke Maeda, Takahiro Ogawa and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
72%Highly rated
?Highly ratedVote to see the score

Imperfect World Models are Exploitable

A novel definition of model exploitation reveals it is essentially unavoidable for large policy sets and cannot be precluded in finite ones, yielding safe planning limits.

Logan M Bhamidipaty, Esmeralda S Whitammer, David Abel, Mykel J Kochenderfer and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 2/5
medium 4/10
strict 2/5
70%Highly rated
?Highly ratedVote to see the score

Multi-Environment POMDPs with Finite-Horizon Objectives

Finite-horizon multi-environment POMDP optimization is PSPACE-complete, and a new practical algorithm significantly outperforms prior methods on benchmarks.

Léonard Brice, Filip Cano, Krishnendu Chatterjee, Thomas Henzinger and 1 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 5 of 20 reviewers recommend it
lenient 1/5
medium 3/10
strict 1/5