Good Papers

Showing papers from UC Berkeley & Amazon Show all papers

45%Niche pick
?Niche pickVote to see the score

When Does Non-Uniform Replay Matter in Reinforcement Learning?

Michal Korniak, Mikołaj Czarnecki, Yarden As, Piotr Miłoś and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

Reward-Conditioned Reinforcement Learning

RCRL conditions agents on reward parameterizations via replay counterfactual rewards, improving sample efficiency and enabling zero-shot behavioral adaptation without extra interaction.

Michal Nauman, Marek Cygan, Pieter Abbeel

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Offline Materials Optimization with CliqueFlowmer

CliqueFlowmer fuses clique-based offline model-based optimization into flow transformers for materials discovery, generating materials that strongly outperform generative baselines.

Jakub Grudzien Kuba, Benjamin K Miller, Sergey Levine, Pieter Abbeel

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 17

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Diffusion Guidance Is a Controllable Policy Improvement Operator

CFGRL links diffusion guidance to policy improvement, training via supervised learning to exceed dataset performance on offline RL without value functions.

Kevin Frans, Seohong Park, Pieter Abbeel, Sergey Levine

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · Code ★ 123

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5