Strategy-Guided Exploration improves LLM agent reinforcement learning by planning in language strategies rather than actions, boosting performance across UI, tool, coding, and embodied tasks.
Private Evolution is recast as learning-augmented clustering to derive tighter bounds via generative models and propose a geometry-aware variant with convergence guarantees.
TGPO decomposes signal temporal logic into timed subgoals and invariant constraints for hierarchical reinforcement learning, achieving 31.6% higher success rates than baselines on complex long-horizon robotics tasks.
Penalized DRO reformulates adversarial risk via optimal transport maps that are cyclically monotone, and enforcing this property via multi-start particle ascent or input-convex networks improves robustness over standard adversarial training.
DLR-Lock replaces MLP weights with deep low-rank residual networks to impose linear backprop memory growth and disrupt optimization, blocking adaptive fine-tuning while preserving model capabilities.
BalCapRL balances RL for MLLM captioning across correctness, coverage, and fluency via normalized multi-objective rewards and length masking, boosting quality metrics substantially.
LensVLM lets VLMs scan compressed rendered text and selectively expand only relevant regions via learned tools, maintaining near-full accuracy at 4.3x compression and outperforming baselines up to 10.1x across text QA benchmarks.
Optimal alignment gains follow a Jeffreys-divergence limit, best-of-N approaches it, and reward hacking grows with error while ensembling mitigates it.