When2Tool finds LLMs linearly encode tool necessity in hidden states, and Probe&Prefill uses this to cut unnecessary tool calls by 48% with minimal accuracy loss.
SimSD proposes a plug-and-play masking strategy that enables token-level speculative decoding in diffusion language models, achieving up to 7.46x faster throughput without training.
AMPS uses instance-aware functional entropy to adaptively steer multimodal model modality preferences, improving control while minimizing inference errors.
Synthetic benchmarks for concept bottleneck models generate controlled labeled datasets to evaluate decision support and automation use cases, diagnose failure modes, and guide testing.
Live Music Diffusion Models modify diffusion inference with block-wise KV caching to surpass discrete autoregressive efficiency, enabling stable alignment via ARC-Forcing and real-time interactive generation on consumer hardware.
Steer2Edit converts inference-time activation steering into training-free, component-level rank-1 weight edits that improve safety, truthfulness, and reasoning efficiency over global interventions.
AnyHand provides 6.6M synthetic RGB-D hand images with occlusions and aligned depth, significantly improving 3D hand pose estimation benchmarks and showing data diversity rivals scale.
Low-rank adaptation regularizes critic learning by constraining updates to low-dimensional subspaces via frozen base weights, reducing loss and improving off-policy RL performance.
LBAC is a programming model that enforces user policies on agentic applications by requiring agents to generate well-typed programs rejected by a type-checker before execution.