When2Tool finds LLMs linearly encode tool necessity in hidden states, and Probe&Prefill uses this to cut unnecessary tool calls by 48% with minimal accuracy loss.
Equilibrium Forcing removes noise conditioning from video diffusion to enable adaptive closed-loop inference that improves generation quality and consistency.
SimSD proposes a plug-and-play masking strategy that enables token-level speculative decoding in diffusion language models, achieving up to 7.46x faster throughput without training.
AMPS uses instance-aware functional entropy to adaptively steer multimodal model modality preferences, improving control while minimizing inference errors.
A framework decomposes RL advantage functions into gradient mass axes, showing trade-offs shift during training and motivating FADE, which adapts weights dynamically to accelerate convergence and improve accuracy-diversity trade-offs.
Synthetic benchmarks for concept bottleneck models generate controlled labeled datasets to evaluate decision support and automation use cases, diagnose failure modes, and guide testing.
INFUSER co-evolves a question generator and solver via influence-guided rewards, improving reasoning by over 20% on math benchmarks without curated data.
ICWBench reveals current LLMs fail at in-context watermarking, and self-distillation with reinforcement learning raises watermark detectability near perfect while preserving quality.
Live Music Diffusion Models modify diffusion inference with block-wise KV caching to surpass discrete autoregressive efficiency, enabling stable alignment via ARC-Forcing and real-time interactive generation on consumer hardware.
By recasting single-shot HDR reconstruction as video diffusion of exposure brackets fused by a lightweight UNet, it achieves high input fidelity and outperforms generative baselines on benchmarks and in 72% of human comparisons.
SWAP assigns step-level length penalties based on reasoning contribution, cutting chain-of-thought length by 64.3% and boosting accuracy 5.7% over base models.
OASIS stabilizes dual-normalized attention-residual architectures via null routing and token-to-depth null coupling, reducing activation outliers by 81.75% and improving low-bit quantized reasoning by 42.11%.
GORMPO integrates generative density estimation into model-based offline RL to restrict policy updates to high-density dataset regions, improving performance by 17% on medical data while linking OOD detection quality to policy gains under stable dynamics.
Steer2Edit converts inference-time activation steering into training-free, component-level rank-1 weight edits that improve safety, truthfulness, and reasoning efficiency over global interventions.
The paper defines meta-design for resource allocation by optimizing upstream design parameters like data, capacity, and quality, and demonstrates the framework in German employment and Ethiopian cash transfer programs.
Direct corpus interaction uses terminal tools to search raw corpora directly, bypassing fixed retrieval interfaces and substantially outperforming sparse, dense, and reranking baselines on agentic search benchmarks.
Calibration with Semantic Reward improves LLM calibration by replacing token-level confidence with direct semantic-space rewards, reducing ECE by up to 40% and raising AUROC by up to 31%.
LatentUM unifies modalities in a shared latent space to enable efficient interleaved cross-modal reasoning and generation, achieving state-of-the-art visual planning and self-reflective generation results.
AnyHand provides 6.6M synthetic RGB-D hand images with occlusions and aligned depth, significantly improving 3D hand pose estimation benchmarks and showing data diversity rivals scale.
Extracting search trees from LLM reasoning traces reveals myopic planning where performance depends on breadth rather than depth, unlike human planning.
ThinkJEPA combines dense JEPA dynamics with sparse VLM reasoning via dual pathways to improve long-horizon latent world modeling and trajectory prediction.
MoveBench introduces a 2.6M-location wildlife movement forecasting benchmark across 110 species and finds existing methods generalize poorly to unseen individuals and deep learning does not consistently beat simpler baselines.
PACEvolve++ adapts evolutionary search policies at test time via advisor-model reinforcement learning, using phase-adaptive optimization to outperform frontier-model baselines across engineering and protein tasks.
Low-rank adaptation regularizes critic learning by constraining updates to low-dimensional subspaces via frozen base weights, reducing loss and improving off-policy RL performance.
SkillRL evolves agents via recursive skill-augmented reinforcement learning with automatic skill discovery and hierarchical library co-evolution, cutting token use while achieving state-of-the-art results across complex tasks.
TIER derives dense tool-use rewards from execution and schemas rather than reference paths, enabling over 90% accuracy on multi-step composition where trajectory supervision fails.
Gradient descent for logistic regression with Gaussian design achieves O(√(‖θ*‖₂⁵d/n)) ℓ₂ error with linear convergence, and a sharper Θ(√(‖θ*‖₂d/n)) rate is tight in high dimensions.
LBAC is a programming model that enforces user policies on agentic applications by requiring agents to generate well-typed programs rejected by a type-checker before execution.