Z-Image

Efficient Image Generation with Single-Stream Diffusion

An efficient 6-billion-parameter foundation model for image generation. Through systematic optimization, it proves that top-tier performance is achievable without relying on enormous model sizes.

欢迎来到 Z-Image 项目!我们很高兴地推出 Z-Image,一个高效的60亿参数图像生成基础模型。它通过系统性的优化证明了顶尖性能的实现无需依赖巨大规模,在照片级真实感图像生成和中英双语文本渲染方面效果突出。

Core Features 核心特性

📸
Photorealistic
Photography-level Realism
摄影级真实感
1 Second
Ultra-fast Inference
超快推理速度
🧠
6B+ Params
Efficient Architecture
60亿高效参数
🔤
Bilingual Text
CN/EN Text Rendering
中英文本渲染
💾
16 GB VRAM
Consumer GPU Friendly
消费级显卡友好
🌍
World Knowledge
Deep Understanding
广博世界知识

Models 模型系列

At just 6 billion parameters, Z-Image produces photorealistic images on par with models an order of magnitude larger.
仅以60亿参数的规模,Z-Image 能生成与参数量大一个数量级的模型相媲美的照片级真实感图像。

🔧 Z-Image-Omni-Base

A foundation model designed for easy fine-tuning, unifying image generation and editing capabilities.

一个易于微调的全能基础模型,统一了图像生成与编辑两大核心功能。

⚡ Z-Image-Turbo

A distilled version with strong photorealistic generation, accurate bilingual text rendering. Achieves top performance with only 8 steps.

蒸馏版本,仅需8步推理即可达到或超越主流竞品性能。

✏️ Z-Image-Edit

Specialized for image editing. Excels at complex instructions, from local modifications to global style transformations.

图像编辑专用版本,精于遵循复杂指令,从局部修改到全局风格变换。

Architecture 架构

Z-Image adopts a Single-Stream Diffusion Transformer architecture. This design unifies the processing of various conditional inputs (text, image embeddings) with noisy image latents into a single sequence fed into the Transformer backbone.

Z-Image 模型采用单流扩散 Transformer 架构。该设计将文本、图像嵌入等多种条件输入与带噪声的图像潜变量统一为单个序列,并送入 Transformer 主干网络进行处理。
Z-Image Architecture: Single-Stream Diffusion Transformer

Arena Ranking 竞技场排名

According to Elo-based Human Preference Evaluation (Alibaba AI Arena), Z-Image shows highly competitive performance against leading models and was presented as one of the strongest open-source image systems of its release window.

根据 Elo 人类偏好评估(在 Alibaba AI Arena 上),Z-Image 与其他领先模型相比表现出极强的竞争力,并在其发布阶段被视为最强的开源图像模型之一。
🏆 Top Open-Source Tier ⚡ 8-Step Inference 💾 16GB VRAM
Z-Image Arena Ranking

1. Efficient Photorealistic Quality 极致高效的照片级真实感

Z-Image-Turbo excels at producing images with photography-level realism, demonstrating fine control over details, lighting, and textures. It balances high fidelity with strong aesthetic quality in composition and overall mood.

Z-Image-Turbo 擅长生成具有摄影级别真实感的图像,能够精细控制画面的细节、光影和纹理。它在保证高保真度的同时,兼顾了构图与整体氛围上的美学表现。

2. Excellent Bilingual Text Rendering 卓越的中英双语文本渲染

Z-Image-Turbo can accurately render Chinese and English text while preserving facial realism and aesthetic composition, with results comparable to top-tier closed-source models. Even in challenging scenarios with small font sizes, it delivers designs that are both textually precise and visually compelling.

Z-Image-Turbo 能准确渲染中英文文本,同时保持人脸真实性和画面美感,效果媲美顶尖闭源模型。即使在小字号等高难度场景下,模型也能高质量地渲染文字。

3. Rich World Knowledge 广博的知识与文化理解

Z-Image possesses a vast understanding of world knowledge and diverse cultural concepts. This allows it to accurately generate famous landmarks, well-known characters, and specific real-world objects.

Z-Image 具备广博的世界知识与对多元文化的深刻理解。这使其能够精确生成各种主题,包括著名地标、知名人物和特定的现实世界物体。

4. Deep Semantic Understanding 引入先验知识的深度语义理解

The powerful prompt enhancer (PE) uses a structured reasoning chain to inject logic and common sense, enabling the model to handle complex tasks like the "chicken-and-rabbit problem" or visualizing classical Chinese poetry.

强大的提示词增强器(PE)通过结构化推理链注入逻辑与常识,使模型能处理诸如"鸡兔同笼"或古诗可视化等复杂任务。在编辑任务中,即使用户指令模糊不清,模型也能运用其推理能力来推断用户的潜在意图。

5. Strong Instruction-Following 强大的指令遵循与创意编辑

Z-Image-Edit can precisely execute complex instructions, such as simultaneously modifying a character's expression and pose while adding specified text. It maintains strong consistency even during significant transformations.

Z-Image-Edit 可精准执行复合指令,如同时修改人物的表情、动作并添加指定文字。即便在如此大幅的图像变换中,它也能保持极高的一致性,体现了对画面每一处元素的精细控制。

Summary 总结

Z-Image demonstrates that top-tier performance is not solely dependent on massive models and computational resources. This lowers the technical and cost barriers for researchers and developers, paving the way for more accessible and innovative applications.

Z-Image 证明了顶尖性能并非只依赖于巨大的模型和计算资源。这为更广泛的研究者和开发者降低了技术与成本门槛,为更多普惠、创新的应用铺平了道路。我们期待社区的积极参与和反馈。

Citation 引用

Welcome to cite our work. 欢迎引用我们的工作。

@misc{z-image-2025,
  title={Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer},
  author={Tongyi Lab},
  year={2025},
  publisher={GitHub},
  journal={GitHub repository},
  howpublished={\url{https://github.com/Tongyi-MAI/Z-Image}}
}
← Back to Quickstart