Good Papers

Showing papers from Montreal Institute for Learning Algorithms, University of Montreal, Université de Montréal Show all papers

57%Worth a look
?Worth a lookVote to see the score

What Kind of Diffusion Models Do We Need in Online Reinforcement Learning?

Zihao Wu, Hongyao Tang, Yi Ma, YAN ZHENG and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Beyond Correctness: Robustness-Driven Evolutionary Self-Training for Large Language Models

Wei Guo, Hongyao Tang, Yi Ma, Jinyi Liu and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Reformulate LLM Reinforcement Learning for Stable Training under Black-box Discrepancy

Jiashun Liu, Runze Liu, Xu Wan, Jing Liang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Does Your Large Language Model Have An Intuitive Sense of The Difficulty of A Question?

Ruitao Wang, Jinyi Liu, Hongyao Tang, Rong Cheng and 8 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Boosting Multiagent Reinforcement Learning at High Replay Ratios with Ensemble Reset

Yaodong Yang, Hongyao Tang, Guangyong Chen, Pheng-Ann Heng

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

LLM RL suffers from training-inference policy mismatch; MIPU optimizes monotonic inference policy improvement to stabilize training and boost reasoning performance.

Jing Liang, Hongyao Tang, Yi Ma, Yancheng He and 7 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 51 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Layerwise LQR for Geometry-Aware Optimization of Deep Networks

Layerwise LQR frames deep network preconditioners as LQR problems to learn scalable structured inverse preconditioners preserving cross-layer geometry, improving optimization dynamics with modest overhead.

Simon Dufort-Labbé, Pierre-Luc Bacon, Razvan Pascanu, Simon Lacoste-Julien and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 2/5
medium 6/10
strict 1/5
72%Highly rated
?Highly ratedVote to see the score

ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching

ForceFlow uses force-aware flow matching with asymmetric multimodal fusion and vision-to-force handover to achieve robust contact-rich manipulation with 37% higher success and stronger zero-shot generalization.

Shuoheng Zhang, Yifu Yuan, Hongyao Tang, YAN ZHENG and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
80%Must read
?Must readVote to see the score

Sparse Layers are Critical to Scaling Looped Language Models

Looped-MoE models scale better than standard transformers via routing divergence that recovers expressivity, and loop boundaries enable efficient early exits with minimal quality loss.

Ryan Lee, Jacob Biloki, Edward J Hu, Jonathan May

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
83%Must read
?Must readVote to see the score

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Embodied-R1.5 is an 8B-parameter embodied foundation model achieving state-of-the-art results on 16 of 24 embodied VLM benchmarks via multi-task balanced RL and a closed-loop planner-grounder-corrector framework.

Yifu Yuan, Yaoting Huang, Xianze Yao, Shuoheng Zhang and 19 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 59 on Hugging Face · Code ★ 60

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 3/5