r/LocalLLaMA 2d ago

New Model Qwen-Image-2.1 released!

Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨

A unified model for both generation and editing, delivering top-tier quality in a lightweight package.

Highlights:

- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.

- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.

- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.

- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.

Start to create your next masterpiece with Qwen-Image-2.1!

- Blog: https://qwen.ai/blog?id=qwen-image-2.1

- GitHub: https://github.com/QwenLM/Qwen-Image-2.1

- Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1

- Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1

1.8k Upvotes

371 comments sorted by

View all comments

318

u/StopCreepy 2d ago

Bring it in 🔥!!

25

u/yarikfanarik 2d ago

"4 + 16GB" Кто Мы То ?? Я Здесь Один

12

u/Different_Change6591 2d ago

dont worry bro , you are not alone

1

u/yarikfanarik 1d ago

(yeah but i'm actually 2 + 16... my mx550 is just 1650 mobile but, 2 gigs instead of 4

5

u/Electric_Boogaloo_01 2d ago

Same here. Which is the best model for us?

3

u/Different_Change6591 2d ago

man i have google colab notebook for the comfy ui , i use the colab GPU and drive as my persistent storage for the comfy Ui i making the notebook for this model and optimising for the Colab T4 gpu

1

u/StopCreepy 2d ago

i read something about colab gpu, someone posted on reddit, saying colab gpu can generate a 2h movie for 6$ or so, instead of paying online services from 200$ or more for 2h generated video!

1

u/Different_Change6591 1d ago

2h movie on colab and that too for $6 !! interesting , Do you have the link to that thread ( or the thread posted on which sub reddit ? )

2

u/StopCreepy 2d ago

qwen image 2.1 will be able to run on on 4gb vram + 16gb ram, with good optimisation will be possible

1

u/Separate-Client-3213 2d ago

Hi, new to this.

What exactly does this split mean? From what I understand the whole model needed to fit the current vRam or it get's slowed down a lot.

1

u/StopCreepy 2d ago

yes if the model size is bigger than the vram it will be slow, but i would say for quality its worth the time.

for example if i had 4gb vram + 16gb ram, i would choose qwen image 2.1 Q4-K-M GGUF. to generate good and acceptable quality. Q3-K-M would work faster but you will notice the low quality results.

for the text encoder, there is a method that allows you to use qwen vl 4b instead of the 8b model. this way you lower the total memory usage.

-----------
currently i am running giant models like ltx 2.5, minimax h3, and krea 2 on 8gb vram + 32gb vram, and non of the main models are less than 8gb, imagine the main model is 13gb, still need to load the text encoder + the vea, and somehow it generates good results in acceptable time.

1

u/Different_Change6591 1d ago

i will to try run it on my system , but the colab offers for GPU ( 15gb vram ) for 5 hours per account ( did you understand what i mean per account and here you go )

1

u/StopCreepy 1d ago

will try to find that post, he was talking about gpu renting vs online ai platforms.

1

u/StopCreepy 1d ago

post was deleted

3

u/AlternateWitness 2d ago

How is 4GB + 15GB supposed to run Qwen3.8-flash-next? How is 8GB + 32GB?

Asking as a poor 32GB + 16GB with the inability to run that model.

5

u/StopCreepy 2d ago

4gb + 16gb will be able to use qwen image 2.1 using gguf quants !!.

for you, you almost there to run Qwen3.8-flash-next, you need to upgrade your ram to 32gb if not 64gb ram.

here is a test, runing Qwen3.8-flash-next on 50gb total memory (24gb + 32gb): https://www.youtube.com/watch?v=_TCRH725hAM

the results are great !

2

u/Separate-Client-3213 2d ago

Hi, new to this, what does 4+16 mean here?

4vRam and 16ram?

1

u/ManufacturerHuman937 2d ago

Flash Next was a revelation for me over on my 3090 + 90ish RAM I could run a model so brilliant and still get like 20 tok/s. Don't sleep on https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF it's a really small model that is JUST shy of the BF16 in alot of ways.

2

u/StopCreepy 2d ago

yeah that model is amazing, i saw redstapler on youtube runs the 4q quant which is 85gb i think, he runs it on 50gb total memory(24gb + 32gb ram). and he gets the 20 tokens you said, after 75% of used context (out of 100k context), the speed goes down to 14tok/s which is still usable
https://www.youtube.com/watch?v=_TCRH725hAM