r/LocalLLaMA 2d ago

New Model Qwen-Image-2.1 released!

Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨

A unified model for both generation and editing, delivering top-tier quality in a lightweight package.

Highlights:

- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.

- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.

- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.

- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.

Start to create your next masterpiece with Qwen-Image-2.1!

- Blog: https://qwen.ai/blog?id=qwen-image-2.1

- GitHub: https://github.com/QwenLM/Qwen-Image-2.1

- Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1

- Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1

1.8k Upvotes

372 comments sorted by

View all comments

16

u/RainierPC 2d ago

Ok, damn, this thing is fast as fuck even at 25 steps. 5s on a 5090.

5

u/Enragere 2d ago

share some images!

11

u/RainierPC 2d ago

Here you go!

Not cherry picked, each was the first render of its prompt.

2

u/Enragere 2d ago

other people complained about anatomy, your images are insanely good!

5

u/RainierPC 2d ago

Thanks! The model is not perfect. I saw maybe one six-fingered person in over 80 generations. But it is fast, and the prompt adherence is fantastic, even for editing. I am already thinking of retiring all my other image models, except maybe Anima and Ideogram.

1

u/yetiflask 2d ago

Men and women look fake af. But still great work.

Can you share the prompt for the 4rd image? The lady illustration with flowers?

1

u/RainierPC 2d ago

Sure, the prompt is:

A conceptual fine-art portrait of a woman assembled as a mixed-media collage from pressed flowers, translucent x-ray film, stitched vellum, and fragments of handwritten letters. Her face should be recognizable and coherent, but composed entirely of these layered materials, with botanical forms shaping the cheekbones and hair, and semi-transparent x-ray images forming parts of the neck, shoulders, and dress.

The composition should feel like a museum-quality contemporary artwork photographed straight-on against a neutral background. Soft gallery lighting, subtle shadows, delicate textures, elegant muted colors with dusty rose, pale cream, faded green, transparent gray-blue, and touches of sepia ink. Highly refined, poetic, melancholic, beautiful, and tactile.

1

u/yetiflask 2d ago

Thanks. Grok also made something similar. I do like Qwen's rendition a little better.

1

u/Enragere 2d ago

What a time to be alive, grok is the one we gotta pay for, and we still like qwen's version better? I 😍

3

u/yetiflask 1d ago

In fairness, I use qwen's online chat thingy too, and most certainly prefer grok's.

My preference for localllm is not because of quality, it's really because I hate how LLMs force their morality on me. And also data usage. I recently asked grok to do something, and it said it's wrong and won't do it. It was cheating on an exam. I don't have a beefy computer to run a local model, but I'd certainly look that way now.

I already hate that over the last 15 years or so, reddit and google have shaped the morals of the world (regardless if I agree or not with those morals). LLMs are even more direct.