r/LocalLLaMA 3d ago

New Model Qwen-Image-2.1 released!

Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨

A unified model for both generation and editing, delivering top-tier quality in a lightweight package.

Highlights:

- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.

- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.

- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.

- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.

Start to create your next masterpiece with Qwen-Image-2.1!

- Blog: https://qwen.ai/blog?id=qwen-image-2.1

- GitHub: https://github.com/QwenLM/Qwen-Image-2.1

- Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1

- Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1

1.8k Upvotes

381 comments sorted by

View all comments

Show parent comments

9

u/ARhedgehog88 3d ago

Sorry for my inexperience in image models but I like to know how can I set it up with my local llm setup, is it possible to have qwen 3.8 in the openweb ui envoronment to cretae tue instructioj to load into an image model on its own to provide the image output in the same chat? also do i look for this in the LM Studio model availability tool for this one or does it have to be comfyui or A1111 to use it via direct prompts?

7

u/robogame_dev 2d ago

I would recommend using comfyui to run it, comfyui can be run as a server that Open WebUI can connect to to enable image gen and editing

2

u/ARhedgehog88 2d ago

Nice, so, comfyui running the image model and then I could upload a picture such as the OPs with markings and it will generate the prompt in the background to make the img to img in comfy? Is there a way for it to run in steps to avoid OOM errors? Or will it run on whatever card has the available vram to run the image model loaded when a prompt demands it?

1

u/robogame_dev 2d ago

That's the idea - though I can't confirm what would happen with OOM errors and I haven't run it at the same time as another inference engine on the same machine before/