Good Papers

Showing papers from University of Montreal, Mila Show all papers

83%Must read
?Must readVote to see the score

Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

Predictive representation learning combined with high-capacity value approximation drives scalable multitask RL, with the simple model-free MR.Q outperforming world-model methods across continuous control tasks.

Johan Obando Ceron, Lu Li, Scott Fujimoto, Pierre-Luc Bacon and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
80%Must read
?Must readVote to see the score

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors

Diagnosis-driven tension management adapts online RL to deployment-specific prior validity shifts, rejecting universal benchmarks for flexible, evidence-guided optimization.

Guozheng Ma, Lu Li, Zilin Wang, Pierre-Luc Bacon and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Layerwise LQR for Geometry-Aware Optimization of Deep Networks

Layerwise LQR frames deep network preconditioners as LQR problems to learn scalable structured inverse preconditioners preserving cross-layer geometry, improving optimization dynamics with modest overhead.

Simon Dufort-Labbé, Pierre-Luc Bacon, Razvan Pascanu, Simon Lacoste-Julien and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 2/5
medium 6/10
strict 1/5
72%Highly rated
?Highly ratedVote to see the score

From Static Policies to Adaptive Priors in Offline Reinforcement Learning

Offline RL should prioritize adaptive policy priors preserving improvement capacity during online updates rather than static conservative deployment.

Tianwei Ni, Vineet Jain, Akash Karthikeyan, Pierre-Luc Bacon

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5