Good Papers

TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization

TextReg mitigates prompt distributional overfitting via regularized text-space optimization, improving out-of-distribution accuracy by up to 16.5% over prior methods.

傅卢成, Ye Yu, Yiyang Wang, Yiqiao Jin, Haibo Jin, B. Aditya Prakash, Haohan Wang

Published May 20, 2026▲ 7 on Hugging FaceCodearXiv ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 4/5
medium 10/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
TextReg delivers compelling out-of-distribution gains through regularized text-space optimization that curbs representational bloat, though its three-part mechanism lacks targeted ablations and unverified generalization beyond reasoning tasks.

Abstract

Large language models (LLMs) are highly sensitive to the prompts used to specify task objectives and behavioral constraints. Many recent prompt optimization methods iteratively rewrite prompts using LLM-generated feedback, but the resulting prompts often become longer, accumulate narrow sample-specific rules, and generalize poorly beyond the training distribution. We study this failure mode as prompt distributional overfitting and argue that it reflects a lack of representation control in discrete text-space optimization. We formalize this view through representational inefficiency, a dual-factor measure that decomposes prompt inefficiency into capacity cost and scope narrowness, attributing distributional prompt overfitting to their coupled growth during optimization. We propose TextReg, a regularization framework that realizes a soft-penalty objective through regularized textual gradients, combining Dual-Evidence Gradient Purification, Semantic Edit Regularization, and Regularization-Guided Prompt Update. Across multiple reasoning benchmarks, TextReg substantially improves out-of-distribution (OOD) generalization, with accuracy gains of up to +11.8% over TextGrad and +16.5% over REVOLVE.