Good Papers

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

RLRT reverses teacher signals in self-distillation to reinforce student reasoning tokens on successful rollouts, substantially outperforming exploration baselines across Qwen3 models.

Jeonghye Kim, Jiwon Jeon, Dongsheng Li, Yuqing Yang

Published 2026Sydney Poster Session 6 · Thu, Dec 10, 5:00 PM–8:00 PM local time · Hall 1-4▲ 15 on Hugging FacearXiv ↗OpenReview ↗

78%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel11/20reviewers recommend it
lenient 3/5
medium 7/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model. While this guidance is useful when the student has failed, on successful rollouts, the same mechanism instead overwrites the student's choices and suppresses it's own reasoning. Therefore, we propose reading the original self-distillation signal in reverse: when the student succeeds along a path the teacher would not have predicted, these tokens reflect its self-driven reasoning. Building on this, we propose RLRT (RLVR with Reversed Teacher), which augments GRPO by reinforcing these tokens on correct rollouts. We interpret this as a new form of exploration in RLVR: not uniform diversity, but valuable exploration grounded in the student's own success. Across base, instruction-tuned, and thinking-tuned Qwen3 checkpoints, RLRT substantially outperforms self-distillation and exploration-based baselines, establishing information asymmetry as a new, principled design axis for RLVR.