When2Tool finds LLMs linearly encode tool necessity in hidden states, and Probe&Prefill uses this to cut unnecessary tool calls by 48% with minimal accuracy loss.
DD-Ranking reveals dataset distillation gains come from extra evaluation techniques rather than image quality, proposing fair metrics to assess true synthetic dataset value.
Equilibrium Forcing removes noise conditioning from video diffusion to enable adaptive closed-loop inference that improves generation quality and consistency.
SimSD proposes a plug-and-play masking strategy that enables token-level speculative decoding in diffusion language models, achieving up to 7.46x faster throughput without training.
AdaCodec uses predictive visual codes to send full reference frames only when unpredictable, cutting video MLLM tokens by 7x while improving long-video benchmark scores and reducing latency.
Live Music Diffusion Models modify diffusion inference with block-wise KV caching to surpass discrete autoregressive efficiency, enabling stable alignment via ARC-Forcing and real-time interactive generation on consumer hardware.
GORMPO integrates generative density estimation into model-based offline RL to restrict policy updates to high-density dataset regions, improving performance by 17% on medical data while linking OOD detection quality to policy gains under stable dynamics.
Steer2Edit converts inference-time activation steering into training-free, component-level rank-1 weight edits that improve safety, truthfulness, and reasoning efficiency over global interventions.
Calibration with Semantic Reward improves LLM calibration by replacing token-level confidence with direct semantic-space rewards, reducing ECE by up to 40% and raising AUROC by up to 31%.
World models hallucinate in low-coverage state-action regions, and coverage-aware sampling plus curiosity rewards detect and mitigate it with minimal data.