Formalizing discretion as a dynamic budget problem yields time-dependent override thresholds and shape-dependent spending rates, with homelessness data showing budget-constrained discretionary patterns.
A process-level latent variable model predicts future behavioral strategy from partial cross-task process traces via transferable person-level representations. In PowerWash Simulator it predicts zone planner versus hopper behavior in held-out levels.
Confidence-based verifier-free test-time scaling fails on complex tasks because high initial confidence signals no exploration; consilience selects rollouts by requiring low early but high final confidence, improving reasoning and coding.
G-Zero uses intrinsic predictive-shift rewards in a verifier-free co-evolutionary framework that enables continuous LLM self-improvement across open-ended unverifiable domains without external judges.