Good Papers

Showing papers from Netflix Show all papers

86%Must read
?Must readVote to see the score

TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment

TeacherGRPO aligns teachers to student distributions via reinforcement learning to overcome reasoning distillation's Gap Curse and improves student performance.

Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng and 4 more

Published Aug 20, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Inverse Reinforcement Learning with Just Classification and a Few Regressions

GenPQR reduces inverse reinforcement learning to policy estimation via classification followed by Q-function regression, yielding modular finite-sample guarantees and improved reward recovery across continuous action spaces.

Lars van der Laan, Nathan Kallus, Aurelien Bibaut

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

The Minimax Rate of Second-Order Calibration

Sech perturbation kernels make calibration functions analytic, enabling polynomial regression to estimate second-order calibration error at the minimax optimal rate of tilde O(1/sqrt(n)). This yields the first finite-sample guarantee for second-order Platt scaling and a bucket-free calibration defin

Kamil Ciosek, Banafsheh Rafiee, Sina Ghiassian, Nicolò Felicioni

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 1/5
medium 7/10
strict 3/5