Good Papers

Showing papers from Google Deepmind Show all papers

57%Worth a look
?Worth a lookVote to see the score

PreFT: Prefill-only finetuning for inference efficiency

Andrew Lanpouthakoun, Aryaman Arora, Zhengxuan Wu, Dhruv Pai and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

On the Selectivity of Generative Models in Structure-Based Drug Design

Ella Miray Rajaonson, Jungyoon Lee, William Chau, Alan Aspuru-Guzik and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ActO: Extracting Action Representations from MLLM Embeddings for Video World Models

Runjia Li, Minghao Chen, Junyu Xie, Philip Torr and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

The Era of Agentic Organization: Learning to Organize with Language Models

Zewen Chi, Li Dong, Qingxiu Dong, Yaru Hao and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Capturing LLM Capabilities via Evidence-Calibrated Query Clustering

ECC calibrates semantic embeddings with limited model comparisons to cluster queries by latent capability demands, improving LLM ranking by ~18 points over semantic baselines and aiding query routing.

Fangzhou Wu, Sandeep Silwal, Qiuyi (Richard) Zhang

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Joint Learning of Hierarchical Neural Options and Abstract World Model

AgentOWL jointly learns hierarchical neural options and an abstract world model for sample-efficient skill acquisition, outperforming baselines on object-centric Atari games with fewer samples and stronger generalization.

Top Piriyakulkij, Wolfgang Lehrach, Kevin Ellis, Kevin Murphy

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Dual-Rate Diffusion: Accelerating diffusion models with an interleaved heavy-light network

Dual-Rate Diffusion accelerates diffusion inference by interleaving sparse heavy context encoders with light denoising models, cutting computation 2-4x without quality loss.

Grigory Bartosh, David Ruhe, Emiel Hoogeboom, Jonathan Heek and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
83%Must read
?Must readVote to see the score

URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets

URDF-Anything+ directly generates simulation-ready articulated URDF models from single RGB images via end-to-end autoregressive diffusion, outperforming prior methods in reconstruction quality, joint estimation, and physical executability.

Zhuangzhe Wu, Yue Xin, Chengkai Hou, Minghao Chen and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

LittleLearner: Language Models Under Pedagogically-Controlled Knowledge Exposure

LittleLearner is a 5B-parameter model trained on grade-capped elementary data to study controlled knowledge acquisition and bounded capability growth.

Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

The FACTS Leaderboard benchmarks large language model factuality across multimodal, parametric, search, and grounding tasks via automated judges.

Aileen Cheng, Alon Jacovi, Amir Globerson, Ben Golan and 36 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 1/5
92%Must read
?Must readVote to see the score

Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking

LLM pipeline evaluation variance is underestimated because design choices are ignored, so corrected intervals restore coverage and cut benchmark gaming.

Solomon Messing

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
19/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 19 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 4/5
80%Must read
?Must readVote to see the score

The Design Space of Tri-Modal Masked Diffusion Models

A tri-modal masked diffusion model pretrained from scratch on text, image-text, and audio-text data achieves strong cross-modal generation and introduces an SDE-based batch-size reparameterization.

Louis Bethune, Victor Guilherme Turrisi da Costa, Bruno Mlodozeniec, Pau Rodriguez and 20 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 2/5
71%Highly rated
?Highly ratedVote to see the score

The Power of Second Order Methods for Sequence Preconditioning

Second-order VAW on short ARX models achieves dimension-free regret O(δ⁻⁴ log² T) for marginally stable linear sequence prediction via universal preconditioning and Faber polynomial analysis.

Annie Marsden, Elad Hazan

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 1/5
medium 3/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

MIND: Monge Inception Distance for Generative Models Evaluation

MIND uses sliced Wasserstein distance via sorting to evaluate generative models with 10x better sample efficiency, 100x faster computation, and greater adversarial robustness than FID.

Quentin Berthet, Clement CREPY, Romuald Elie, Klaus Greff and 2 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs

IDEAFix evaluates LLM divergent thinking via controlled design scenarios and defixation prompts, showing task formulation and simple prompting boost originality but homogenization persists.

Florian Carichon, Soumya Sharma, Meaghan J. Girard, Romain Rampa and 1 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5