Good Papers

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Z-Image is a 6B-parameter diffusion image generator that achieves leading open-source performance with only 314K GPU hours, sub-second inference, and consumer-hardware compatibility.

Image Team, Cai, Huanqia, Cao, Sihan, Du, Ruoyi, Gao, Peng, Hao, Aiming, Hoi, Steven, Hou, Zhaohui, Huang, Shijie, Jiang, Dengyang, Jiang, Yuming, Jin, Xin, Li, Liangchen, Li, Zhen, Li, Zhong-Yu, Liu, David, Liu, Dongyang, Wu, Qilong, Yu, Feng, Zhan, Zechao, Zhang, Chi, Zhang, Shifeng, Zhou, Ruikai, Zhou, Shilin

Published Nov 27, 20251 citation▲ 249 on Hugging FaceCode ★ 12,067arXiv ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 5/5
medium 4/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Z-Image earns praise for delivering competitive photorealism and bilingual text at 6B parameters with leaner training costs, yet the discussion remains skeptical of its vague "comparable" benchmarks, missing head-to-head metrics against 20B-80B rivals, and whether sub-second…

Abstract

The landscape of high-performance image generation models is currently dominated by proprietary systems, such as Nano Banana Pro and Seedream 4.0. Leading open-source alternatives, including Qwen-Image, Hunyuan-Image-3.0 and FLUX.2, are characterized by massive parameter counts (20B to 80B), making them impractical for inference, and fine-tuning on consumer-grade hardware. To address this gap, we propose Z-Image, an efficient 6B-parameter foundation generative model built upon a Scalable Single-Stream Diffusion Transformer (S3-DiT) architecture that challenges the "scale-at-all-costs" paradigm. By systematically optimizing the entire model lifecycle -- from a curated data infrastructure to a streamlined training curriculum -- we complete the full training workflow in just 314K H800 GPU hours (approx. $630K). Our few-step distillation scheme with reward post-training further yields Z-Image-Turbo, offering both sub-second inference latency on an enterprise-grade H800 GPU and compatibility with consumer-grade hardware (<16GB VRAM). Additionally, our omni-pre-training paradigm also enables efficient training of Z-Image-Edit, an editing model with impressive instruction-following capabilities. Both qualitative and quantitative experiments demonstrate that our model achieves performance comparable to or surpassing that of leading competitors across various dimensions. Most notably, Z-Image exhibits exceptional capabilities in photorealistic image generation and bilingual text rendering, delivering results that rival top-tier commercial models, thereby demonstrating that state-of-the-art results are achievable with significantly reduced computational overhead. We publicly release our code, weights, and online demo to foster the development of accessible, budget-friendly, yet state-of-the-art generative models.