TACache decomposes rectified flow velocity errors into magnitude and direction components to skip steps and reconstruct velocities without extra evaluations, achieving up to 4.14x faster image and 2.11x faster video generation.
WorldAct converts static generated 3D worlds into editable, interaction-ready scenes via multimodal decomposition and object reconstruction to enable manipulation and embodied tasks.
TurboVGGT enables fast multi-view 3D reconstruction via adaptive alternating attention that balances sparse global and local frame attention while maintaining competitive quality.
Memory-R2 proposes LoGo-GRPO to enable fair credit assignment for memory-augmented LLM agents across long multi-session horizons via local rerollouts and shared-parameter co-learning.
HiFloat4 enables stable FP4 LLM pretraining without stabilization stacks, achieving 1.55% relative loss versus 1.79% for MXFP4 and 2.00% for NVFP4 on Ascend NPUs.
Supervised fine-tuning harms reasoning by memorizing surface correlations rather than theorem application; Theorem-SFT improves MATH and GeoQA scores by teaching explicit rule invocation.
PSD accelerates diffusion LLM inference via adaptive parallel unmasking and multi-depth speculative drafts with hierarchical verification, achieving up to 5.5x tokens per pass with near-greedy accuracy.
BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
MemDLM augments diffusion language model training via bi-level optimization with parametric memory, improving convergence, long-context representations, and needle retrieval.
Imagining in 360° decouples humanoid visual search into an Imaginator predicting spatial priors and an Actor using them to improve search efficiency without costly trajectory annotations.
VETime unifies temporal and visual modalities via fine-grained alignment and dynamic fusion for zero-shot time-series anomaly detection, outperforming state-of-the-art models with lower overhead.
PEARL integrates solvers into an interactive optimization modeling loop to iteratively revise formulations using execution feedback, substantially boosting verified solve rates and enabling a small model to outperform a much larger baseline.
Flexible Context Parallelism adaptively reconfigures communication groups to eliminate load imbalance and redundant communication, achieving up to 1.46x training speedup over Megatron-LM and DeepSpeed.
Multimodal LLM safety failure stems from geometry collapse along refusal directions caused by modality drift, which adaptive drift correction and self-rectification restore without training.
ForceFlow uses force-aware flow matching with asymmetric multimodal fusion and vision-to-force handover to achieve robust contact-rich manipulation with 37% higher success and stronger zero-shot generalization.
FaithSieve decomposes math proofs into local reasoning units and verifies them in Lean with semantic alignment gating, achieving over 81% first-error localization accuracy on expert benchmarks.
TextPro-SLM minimizes speech-text modality gap by feeding prosody-aware text inputs to LLMs, cutting the gap at 3B and 7B scales with only ~1,000 hours of audio.
LAMP adaptively selects key transformer components for high-precision recomputation, cutting inference error by up to two orders of magnitude with minimal overhead.
Linguistic Trajectory Encoding compresses long-horizon object motion into hybrid language-spatial-visual timelines, outperforming baselines on multi-day spatial memory benchmarks with high compression and sub-second queries.
Inline Critic uses learnable tokens to critique frozen image-editing models at intermediate layers, steering hidden states during the forward pass to achieve state-of-the-art results.
Seg3DParts uses segmentation-grounded generation with structured cross-part interaction to produce controllable, coherent part-level 3D meshes from single images while introducing the PartObjectNet dataset.
STILL introduces self-saliency token selection and norm-preserved feature maps to linearize LLMs, matching original performance with up to 86.2% long-context gains.
HiFloat4 enables end-to-end FP4 reinforcement learning by fixing rollout activation underflow with Rollout-ResQ, cutting accuracy gaps to 1.1% versus BF16.
MIRAGE learns continuous latent reasoning for mobile agents, cutting decoded tokens 75% while matching explicit chain-of-thought accuracy and improving baselines up to 10.2 points via generative world modeling.
IRR-Drive uses adaptive multimodal text and BEV reflection to self-correct driving intentions before trajectory generation, achieving state-of-the-art NAVSIM results.
Systematic analysis of triangular inversion for delta-rule linear transformers yields algorithms with up to 4.3x speedup on NPUs and preserved end-to-end accuracy across low-precision settings.
SelfBootTok decomposes image tokens into global and local groups via self-bootstrapped learning to shift detail burden to the tokenizer, cutting generator computation by ~40% while achieving state-of-the-art 1.56 gFID with 64 tokens.
SePO evolves its own prompt agent's system prompt via self-referential open-ended search to optimize task agents, outperforming baselines by 4.49 points across five benchmarks.
G²TR uses generation-branch signals to reduce visual tokens in unified multimodal models, cutting prefill computation by 1.94× while preserving reasoning and editing performance.
PermuQuant reorders diffusion model channels by statistical similarity before per-group quantization to reduce low-bit quantization error and accelerate inference.