Good Papers

Showing papers from DeepMind Show all papers

74%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026DeepMindDeep RL

Delightful Distributed Policy Gradient

Delightful Policy Gradient gates distributed updates with delight (advantage times surprisal) to suppress high-surprisal failures while preserving rare successes, outperforming importance-weighted methods under staleness, bugs, and corruption.

Ian Osband

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

0% Readers0 of 1 upvoted
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 3/5
medium 10/10
strict 3/5
45%Niche pick
?Niche pickVote to see the score

DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents

Yun-Shiuan Chuang, Ruixuan Tu, Chengtao Dai, You Li and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Off-policy Learning with Excursion Policies

Jiamin He, Mark Rowland, Daniel (Zhaohan) Guo, Hado van Hasselt and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Ballad: Bandit-Based LLM Routing for Automated Heuristic Discovery

Samidha Verma, Ankit Anand, Sayan Ranu

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

A Generative Model of Contextual Integrity: Appropriate vs. Inappropriate Information Sharing

Omer Ebead, Juan Formanek, Joel Leibo

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
86%Must read
?Must readVote to see the score

Inverting the Bellman Equation: From $Q$-Values to World Models

Value-based agents trained on diverse reward functions implicitly encode world models, extractable via P-learning, with sufficient conditions for exact dynamics recovery and cross-goal generalization.

Alistair Letcher, Mattie Fellows, Alexander D. Goldie, Jonathan Richens and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 3/5
80%Must read
?Must readVote to see the score

Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts

Persona Generators use evolutionary code optimization to expand brief context descriptions into diverse synthetic populations maximizing opinion and preference coverage. Evolved generators substantially outperform baselines across six diversity metrics by spanning rare trait combinations.

Davide Paglieri, Logan Cross, William Cunningham, Joel Leibo and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

A Unified Framework for Adversary-Aware Differential Privacy Bounds

A unified framework bounds DP privacy leakage against multi-target membership, attribute, and reconstruction attacks using only privacy parameters and adversarial baseline success rates.

Marika Swanberg, Meenatchi Sundaram Muthu Selva Annamalai, Jamie Hayes, Borja Balle and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
89%Must read
?Must readVote to see the score

Scaling Storm-Resolving Atmospheric AI Simulation to the Entire Planet

STRATA is an autoregressive AI emulator for global storm-resolving atmospheric dynamics that achieves 50× better energy efficiency and stable 24-hour rollouts on limited training data.

Zeyuan Hu, Noah Brenowitz, Akshay Subramaniam, Jaideep Pathak and 8 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
74%Highly rated
?Highly ratedVote to see the score

Paradoxes of Game Theoretic Equilibria and Price of Anarchy

Static equilibrium and black-box regret analysis obscure dynamic disequilibrium; worst-case equilibria are unstable saddles, PoA becomes unbounded under affine costs, and discrete-time learning drives chaos with exponentially degrading inefficiency.

Georgios Piliouras, Ian Gemp, Siqi Liu, Luke Marris

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 1/5
medium 6/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

MIND: Monge Inception Distance for Generative Models Evaluation

MIND uses sliced Wasserstein distance via sorting to evaluate generative models with 10x better sample efficiency, 100x faster computation, and greater adversarial robustness than FID.

Quentin Berthet, Clement CREPY, Romuald Elie, Klaus Greff and 2 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
83%Must read
?Must readVote to see the score

Utility-Constrained Policy Optimization

A practical utility-constrained MDP framework enables risk-sensitive safety constraints and flexible limit adjustments without retraining, matching or outperforming Safety Gymnasium baselines.

Mehrdad Moghimi, Bernardo Avila Pires

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Understanding Generalization Requires Universal Induction

Classical statistics cannot justify AI inductive biases, so universal induction relativized to preexisting information formally defines the limits of learnable prediction.

Aram Ebtekar, Marcus Hutter, Danica J. Sutherland

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 0/5