TurboVGGT enables fast multi-view 3D reconstruction via adaptive alternating attention that balances sparse global and local frame attention while maintaining competitive quality.
HiFloat4 enables stable FP4 LLM pretraining without stabilization stacks, achieving 1.55% relative loss versus 1.79% for MXFP4 and 2.00% for NVFP4 on Ascend NPUs.
Multimodal LLM safety failure stems from geometry collapse along refusal directions caused by modality drift, which adaptive drift correction and self-rectification restore without training.
HiFloat4 enables end-to-end FP4 reinforcement learning by fixing rollout activation underflow with Rollout-ResQ, cutting accuracy gaps to 1.1% versus BF16.
Systematic analysis of triangular inversion for delta-rule linear transformers yields algorithms with up to 4.3x speedup on NPUs and preserved end-to-end accuracy across low-precision settings.