Good Papers

Showing papers from Google DeepMind / U. de Montreal / Mila Show all papers

83%Must read
?Must readVote to see the score

Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

Predictive representation learning combined with high-capacity value approximation drives scalable multitask RL, with the simple model-free MR.Q outperforming world-model methods across continuous control tasks.

Johan Obando Ceron, Lu Li, Scott Fujimoto, Pierre-Luc Bacon and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

A Mechanistic Analysis of Looped Reasoning Language Models

Looped reasoning models converge to cyclic fixed points that stabilize attention and repeat feedforward inference stages iteratively, with recurrence size and normalization affecting stability.

Hugh Blayney, Alvaro Arroyo, Johan Obando Ceron, Pablo Samuel Castro and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 0/5
88%Must read
?Must readVote to see the score

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

Agentick unifies RL and foundation model agent evaluation across 37 tasks, finding no dominant approach and substantial room for improvement.

Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 4/5
86%Must read
?Must readVote to see the score

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions

Gaussian trust region reshaping replaces monotonic divergence penalties with bounded non-monotonic constraints, unlocking efficient behavior transitions in non-stationary reinforcement learning.

Bingxu Liu, Jiashun Liu, Johan Obando Ceron, Hao Wang and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5