Strategy-Guided Exploration improves LLM agent reinforcement learning by planning in language strategies rather than actions, boosting performance across UI, tool, coding, and embodied tasks.
Private Evolution is recast as learning-augmented clustering to derive tighter bounds via generative models and propose a geometry-aware variant with convergence guarantees.
TGPO decomposes signal temporal logic into timed subgoals and invariant constraints for hierarchical reinforcement learning, achieving 31.6% higher success rates than baselines on complex long-horizon robotics tasks.
Penalized DRO reformulates adversarial risk via optimal transport maps that are cyclically monotone, and enforcing this property via multi-start particle ascent or input-convex networks improves robustness over standard adversarial training.
DLR-Lock replaces MLP weights with deep low-rank residual networks to impose linear backprop memory growth and disrupt optimization, blocking adaptive fine-tuning while preserving model capabilities.
BalCapRL balances RL for MLLM captioning across correctness, coverage, and fluency via normalized multi-objective rewards and length masking, boosting quality metrics substantially.
LensVLM lets VLMs scan compressed rendered text and selectively expand only relevant regions via learned tools, maintaining near-full accuracy at 4.3x compression and outperforming baselines up to 10.1x across text QA benchmarks.
Optimal alignment gains follow a Jeffreys-divergence limit, best-of-N approaches it, and reward hacking grows with error while ensembling mitigates it.
CM2 replaces verifiable outcome rewards with checklist rewards for multi-turn tool-use RL, improving 8B models by 8, 12 points on agent benchmarks using simulated environments.
AME-TS guides sparse mixture-of-experts routing via temporal structure descriptors to improve forecasting accuracy and specialization stability with fewer activated parameters.
STARFlow2 unifies multimodal generation by vertically interleaving a pretrained vision-language model with an autoregressive normalizing flow under shared causal masking, enabling cache-friendly interleaved text-image generation with strong benchmark performance.
Algorithms preserve accurate subpopulation classification rates and enable loss minimization, but simultaneously achieving both is computationally infeasible despite Bayes-optimal feasibility.
TIDE injects token identity into every layer via EmbeddingMemory to fix rare-token undertraining and contextual collapse, improving language modeling and downstream performance.
DACA-GRPO improves diffusion language model RL via denoising progress scores and stratified masking likelihood, boosting math, code, and constraint benchmarks by up to 36.3pp.
HACRL enables heterogeneous agents to share verified rollouts during collaborative on-policy training and execute independently at inference, with HACPO improving all agents by 3.6% over baselines at half the rollout cost.
A tri-modal masked diffusion model pretrained from scratch on text, image-text, and audio-text data achieves strong cross-modal generation and introduces an SDE-based batch-size reparameterization.
Action Images formulates robot policy learning as multiview video generation using interpretable pixel-grounded action images, enabling zero-shot control without separate policy heads and improving video-action joint generation.
SkipSR accelerates diffusion video super-resolution by skipping low-detail regions identified in low-resolution inputs, achieving up to 60% faster latency without quality loss.
Constant-context skill learning embeds recurring agent workflows into lightweight modules via step-level SFT and online RL, cutting prompt tokens 2-7x while matching state-of-the-art success on ALFWorld, WebShop, and SciWorld.
Trajectory-Shaped Discrete Flow Matching guides discrete flow matching training via an energy-based midpoint evaluator, letting small students outperform large teachers with 32% lower perplexity at 128x speed.