Good Papers

Showing papers from Google & Google DeepMind Show all papers

86%Must read
?Must readVote to see the score

Bayesian Preference Learning for Test-Time Steerable Reward Models

ICRM enables test-time steerable reward models via Bayesian variational inference over preferences, improving multi-objective alignment, calibration, and math reasoning.

Jiwoo Hong, Shao Tang, Zhipeng Wang

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 2/5