We are excited to introduce LLaDA-Image and LLaDA-Image-Turbo.
Highlights:
π¨ High-quality generation for photorealistic images, posters, ads, and bilingual typography
πͺ One 6B DiT unifies text-to-image generation and instruction-guided editing
Both backone and Image-Gen are diffusion models, trained in a unified framework.
β‘ Fast 2-4-step inference with LLaDA-Image-Turbo
4
20
150
9,213
High-quality text-to-image pre-training need not begin with image-text pairs.
πLLaDA-Image builds its visual generative prior from images. Of ~220M cumulative generation-training samples, >90% use image-only supervision; paired data handles later language alignment.
Image-only data builds the visual world; image-text pairs connect language to it.
Both backone and Image-Gen are diffusion models, trained in a unified framework.
1
6
1,306
On Qwen-Image-Bench, LLaDA-Image scores 53.53 on the English track and 53.38 on the Chinese track, advance rankings on both among the open-source models listed in the technical report.
Generation and editing share one backbone. The same model family offers two deployment profiles: Base prioritizes full quality, while Turbo prioritizes inference speed.
#OpenSourceAI #dLLM #inclusionAI
Try out and explore LLaDA-Image, the open release includes Base and Turbo weights, training and inference code, and the complete training recipes:
π€ Hugging Face: huggingface.co/collections/iβ¦
π» Code: github.com/inclusionAI/LLaDAβ¦
π Technical Report: arxiv.org/pdf/2609.03796
Sep 4, 2026 Β· 2:17 AM UTC
1
10
962



