VideoRLVR applies reinforcement learning with verifiable rewards to video diffusion models, improving rule-consistent visual reasoning and cutting training latency 40% via early-step optimization.
TTCD uses a long-window teacher to supervise a short-window student's fast weights via hidden-state discrepancy, allocating limited memory to future-relevant context and outperforming existing long-context methods with minimal architectural changes.
Learning-augmented online scheduling achieves O(1)-competitive latency with O(1) preemptions per job on parallel machines, with overhead scaling logarithmically in prediction error.