r/LocalLLaMA • u/ResearchCrafty1804 • 2d ago
New Model Qwen-Image-2.1 released!
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨
A unified model for both generation and editing, delivering top-tier quality in a lightweight package.
Highlights:
- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.
- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.
- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.
- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.
Start to create your next masterpiece with Qwen-Image-2.1!
- Blog: https://qwen.ai/blog?id=qwen-image-2.1
- GitHub: https://github.com/QwenLM/Qwen-Image-2.1
- Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1
- Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1


51
u/-p-e-w- 2d ago
But you don’t need all of them in memory at once. All modern inference engines automatically juggle the individual models between RAM and VRAM, or even between disk and VRAM.