Good Papers

G-Zero: Self-Play for Open-Ended Generation from Zero Data

G-Zero uses intrinsic predictive-shift rewards in a verifier-free co-evolutionary framework that enables continuous LLM self-improvement across open-ended unverifiable domains without external judges.

Chengsong Huang, Haolin Liu, Tong Zheng, Runpeng(Leo) Dai, Langlin Huang, JINYUAN LI, Zongxia Li, Zhepei Wei, Yu Meng, Jiaxin Huang

Published 2026Atlanta Poster Session 2 · Wed, Dec 9, 4:30 PM–7:30 PM local time · Hall C1▲ 17 on Hugging FaceCode ★ 30arXiv ↗OpenReview ↗

72%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel8/20reviewers recommend it
lenient 3/5
medium 5/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Self-evolving LLMs excel in verifiable domains but struggle in open-ended tasks, where reliance on proxy LLM judges introduces capability bottlenecks and reward hacking. To overcome this, we introduce G-Zero, a verifier-free, co-evolutionary framework for autonomous self-improvement. Our core innovation is Hint-$δ$, an intrinsic reward that quantifies the predictive shift between a Generator model's unassisted response and its response conditioned on a self-generated hint. Using this signal, a Proposer model is trained via GRPO to continuously target the Generator's blind spots by synthesizing challenging queries and informative hints. The Generator is concurrently optimized via DPO to internalize these hint-guided improvements. Theoretically, we prove a best-iterate suboptimality guarantee for an idealized standard-DPO version of G-Zero, provided that the Proposer induces sufficient exploration coverage and the data filteration keeps pseudo-label score noise low. By deriving supervision entirely from internal distributional dynamics, G-Zero bypasses the capability ceilings of external judges, providing a scalable, robust pathway for continuous LLM self-evolution across unverifiable domains.