86%Must read
?Must readVote to see the score

Bayesian Preference Learning for Test-Time Steerable Reward Models
ICRM enables test-time steerable reward models via Bayesian variational inference over preferences, improving multi-objective alignment, calibration, and math reasoning.
Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026
– ReadersNo votes yet
14/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 2/5