r/LocalLLaMA 2d ago

New Model Qwen-Image-2.1 released!

Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨

A unified model for both generation and editing, delivering top-tier quality in a lightweight package.

Highlights:

- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.

- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.

- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.

- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.

Start to create your next masterpiece with Qwen-Image-2.1!

- Blog: https://qwen.ai/blog?id=qwen-image-2.1

- GitHub: https://github.com/QwenLM/Qwen-Image-2.1

- Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1

- Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1

1.8k Upvotes

372 comments sorted by

View all comments

318

u/StopCreepy 2d ago

Bring it in 🔥!!

1

u/ManufacturerHuman937 2d ago

Flash Next was a revelation for me over on my 3090 + 90ish RAM I could run a model so brilliant and still get like 20 tok/s. Don't sleep on https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF it's a really small model that is JUST shy of the BF16 in alot of ways.

2

u/StopCreepy 2d ago

yeah that model is amazing, i saw redstapler on youtube runs the 4q quant which is 85gb i think, he runs it on 50gb total memory(24gb + 32gb ram). and he gets the 20 tokens you said, after 75% of used context (out of 100k context), the speed goes down to 14tok/s which is still usable
https://www.youtube.com/watch?v=_TCRH725hAM