TACache decomposes rectified flow velocity errors into magnitude and direction components to skip steps and reconstruct velocities without extra evaluations, achieving up to 4.14x faster image and 2.11x faster video generation.
WorldAct converts static generated 3D worlds into editable, interaction-ready scenes via multimodal decomposition and object reconstruction to enable manipulation and embodied tasks.
TurboVGGT enables fast multi-view 3D reconstruction via adaptive alternating attention that balances sparse global and local frame attention while maintaining competitive quality.
Memory-R2 proposes LoGo-GRPO to enable fair credit assignment for memory-augmented LLM agents across long multi-session horizons via local rerollouts and shared-parameter co-learning.
HiFloat4 enables stable FP4 LLM pretraining without stabilization stacks, achieving 1.55% relative loss versus 1.79% for MXFP4 and 2.00% for NVFP4 on Ascend NPUs.
Supervised fine-tuning harms reasoning by memorizing surface correlations rather than theorem application; Theorem-SFT improves MATH and GeoQA scores by teaching explicit rule invocation.
PSD accelerates diffusion LLM inference via adaptive parallel unmasking and multi-depth speculative drafts with hierarchical verification, achieving up to 5.5x tokens per pass with near-greedy accuracy.
BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
MemDLM augments diffusion language model training via bi-level optimization with parametric memory, improving convergence, long-context representations, and needle retrieval.
Imagining in 360° decouples humanoid visual search into an Imaginator predicting spatial priors and an Actor using them to improve search efficiency without costly trajectory annotations.
VETime unifies temporal and visual modalities via fine-grained alignment and dynamic fusion for zero-shot time-series anomaly detection, outperforming state-of-the-art models with lower overhead.
PEARL integrates solvers into an interactive optimization modeling loop to iteratively revise formulations using execution feedback, substantially boosting verified solve rates and enabling a small model to outperform a much larger baseline.
Flexible Context Parallelism adaptively reconfigures communication groups to eliminate load imbalance and redundant communication, achieving up to 1.46x training speedup over Megatron-LM and DeepSpeed.
Multimodal LLM safety failure stems from geometry collapse along refusal directions caused by modality drift, which adaptive drift correction and self-rectification restore without training.
ForceFlow uses force-aware flow matching with asymmetric multimodal fusion and vision-to-force handover to achieve robust contact-rich manipulation with 37% higher success and stronger zero-shot generalization.