Good Papers

HiFloat4 Format for Language Model Pre-training on Ascend NPUs

HiFloat4 enables stable FP4 LLM pretraining without stabilization stacks, achieving 1.55% relative loss versus 1.79% for MXFP4 and 2.00% for NVFP4 on Ascend NPUs.

Mehran Taghian Jazi, Yunke Peng, Xing Huang, Yao Wang, Yaoyuan Wang, Wei Guo, Yuanyong Luo, Tianchi Hu, JUNSONG WANG, Xin Wang, Hu Liu, Yu Cheng, Yu Z Wei, Hongliang Li, Mehdi Rahimifar, Lei YAN, wxuefei, Zhuang Ma, Liulei, Hui Yu, Anandharaju D Raju, Hoang Le, Hei Yi Mak, Tanzila Rahman, Shadan Golestan

Published 2026Sydney Poster Session 4 · Wed, Dec 9, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

88%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel15/20reviewers recommend it
lenient 5/5
medium 7/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Training large foundation models at low numerical precision is one of the most promising directions for reducing the compute and memory cost of modern AI. Recent 4-bit floating-point formats such as MXFP4 and NVFP4 can be applied to linear GEMM operations in LLMs, but their limited dynamic range introduces numerical instability that prior work addresses by stacking stabilization mechanisms, typically executed at higher precision and partially eroding the efficiency gains that motivate FP4. In this work, we argue that numerical format design is itself a first-class lever for stable FP4 training, and present the first systematic study of FP4 LLM pretraining on energy-efficient Huawei Ascend NPUs. We compare the recently proposed HiFloat4 (HiF4) format against both MXFP4 and NVFP4, holding one recipe fixed across all three formats across dense (OpenPangu-1B, Llama3-8B) and Mixture-of-Experts (Qwen3-MoE-30B) architectures and executing all linear and expert GEMMs in FP4. At matched storage --- NVFP4 and HiF4 both spend 4.5 bits per value --- the three formats differ far more in what they require before they will train at all than in final accuracy. HiF4 reaches a relative loss of 1.55\% with no stabilization, below fully stabilized MXFP4 (1.79\%) and below NVFP4 carrying the per-tensor scaling it cannot train without (2.00\%); NVFP4 diverges under every combination of stochastic rounding and Hadamard transform we tried. Our results suggest that stable, accurate FP4 training does not require an ever-growing stack of stabilization techniques; it requires the right numerical format.