Attention-based sampling orders diffusion language model tokens by attention-matrix column sums to maximize likelihood and improve generation quality with greater parallelism.
HA-HOI reconstructs physically plausible 4D human-object interactions from monocular video by anchoring object motion to human action and refining via physics simulation. It improves alignment, contact consistency, and simulation readiness over prior monocular reconstruction methods.
PEARL integrates solvers into an interactive optimization modeling loop to iteratively revise formulations using execution feedback, substantially boosting verified solve rates and enabling a small model to outperform a much larger baseline.
AHPA adaptively selects hierarchical VAE feature priors via a timestep-conditioned router to match diffusion transformer alignment granularity to denoising needs, improving convergence without inference overhead.
SkillMigrator learns reusable web skills via transferable interaction patterns matched by layout similarity to reduce LLM actions 8-10% across WebArena and Mind2Web.
PRECISE introduces an SDE-consistent stochastic sampler balancing exploration and stability for RL post-training of flow-matching models, enabling faster, more stable reward optimization with significantly reduced training time.