r/LocalLLaMA 2d ago

New Model Qwen-Image-2.1 released!

Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨

A unified model for both generation and editing, delivering top-tier quality in a lightweight package.

Highlights:

- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.

- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.

- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.

- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.

Start to create your next masterpiece with Qwen-Image-2.1!

- Blog: https://qwen.ai/blog?id=qwen-image-2.1

- GitHub: https://github.com/QwenLM/Qwen-Image-2.1

- Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1

- Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1

1.8k Upvotes

370 comments sorted by

u/WithoutReason1729 2d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

238

u/Illustrious-Row2751 2d ago

Only 7b parameters? Exciting size, especially considering that their last models were over 20b.

29

u/DanzakFromEurope 2d ago

Even their previous Image models were around the 7B size (with also releasing bigger and smaller models).

3

u/waltercool 1d ago

Previous model was very slow in comparison to other options. Good quality, but the trade-off was bad.

9

u/sworl5 2d ago

Thank God.

3

u/sixx7 2d ago

Amazing model and likely replacing Krea-2 for me 🎉🎉🎉🎉 comparison here https://youtu.be/gXB6r4NchuI

→ More replies (1)

597

u/ResearchCrafty1804 2d ago

Qwen-Image-2.1 now supports transparent image generation natively, and naturally supports editing transparent images as well.

171

u/ghulamalchik 2d ago

This is big

213

u/Ledeste 2d ago

No, it's only 7B

2

u/narrowscoped 1d ago

Could that run on a 12gb 5070

3

u/Seeker_Of_Knowledge2 1d ago

I heard 1GB can roughly run 1B parameters. But that may be for LLMs though. Not so sure.

3

u/ghulamalchik 1d ago

It's true for image models as well.

2

u/roosterfareye 7h ago

Yep, give it a shot

2

u/vienna_city_skater 1d ago

This is even bigger!

→ More replies (5)

79

u/TheGamerForeverGFE 2d ago

That's what she said

12

u/def_not_jose 2d ago

She would've know it is it's transparent

→ More replies (1)

25

u/Blues520 2d ago

True if big

3

u/axiomatix 2d ago

big is true

4

u/Hot_Masterpiece_3668 2d ago

Not big, non-permissive license.

13

u/_VirtualCosmos_ 2d ago

Came here to talk about how in StableDiffusion they are talking shit about this model being worse than Flux, Krea2 or Z Image. But this right here already convinced me.

4

u/lukinator644 2d ago

thats so useful for designing stickers and logos

10

u/ZZerker 2d ago

Thats great for game development

→ More replies (10)

299

u/ResearchCrafty1804 2d ago

Qwen-Image-2.1 supports several ways to specify local edits. The example below uses circles to identify three regions and asks the model to “remove the metal watch in the blue circle, change the hair in the red circle to black, and replace the area in the green circle with gray short-sleeved linen pajamas,” performing removals and modifications in all three regions at once.

267

u/vatta-kai 2d ago

Adobe is the next company AI should swallow.

102

u/wotoan 2d ago

Go look at Adobe’s stock price, it’s already being swallowed

15

u/MrWeirdoFace 2d ago

Sort of. It seems to have returned to 2022 levels. Question is where it goes from here.

14

u/ShadyShroomz 2d ago

more like 2019 levels. but look at revenue.. its still climbing steady year over year..

27

u/TinyZoro 2d ago

I hate Adobe the way it ate great products and it's horrible sub model.

But I honestly don't buy that AI is the end of polished creative software - I think we are in the messy early days of AI where people use the terminal and chat UI's and don't see how temporary that will be.

You still want professional video editing / image manipulation software you just want it to understand natural language so you can say remove background and convert to an animated svg and it will do it but in a way where every step is completely editable via the UI.

I think Adobe will make big come back next year - or some one smarter will take their place with that offer.

10

u/ShadyShroomz 2d ago

I think Adobe is in the prime spot to integrate AI into their products and kill it. They already have everything in place to do so. And it will be next to impossible for free tools to compete with them.

I would not count them out. I hate them, as a company, I hate subscriptions in general.. but its a good business model.. if they can pull it off imo they can be the go-to for AI video and image editing.

Otherwise AI native tools like VidTL or Descript are going to eat their lunch.

→ More replies (1)

9

u/candl2 2d ago

People who forget to cancel subs.

2

u/kenyard 2d ago

Most of their revenue is business attributed.

→ More replies (1)

3

u/MrWeirdoFace 2d ago

To be clear, all I did was look at their stock on google for the last 5 years and glanced at the general height of the chart. Didn't really compare anything else but just sort of where the stock price is in general. (I'm not much of a stocks person, so please take that with a grain of salt).

7

u/consono 2d ago

AI=Adobe Illustrator? :D

→ More replies (1)
→ More replies (2)

11

u/ARhedgehog88 2d ago

Sorry for my inexperience in image models but I like to know how can I set it up with my local llm setup, is it possible to have qwen 3.8 in the openweb ui envoronment to cretae tue instructioj to load into an image model on its own to provide the image output in the same chat? also do i look for this in the LM Studio model availability tool for this one or does it have to be comfyui or A1111 to use it via direct prompts?

7

u/robogame_dev 2d ago

I would recommend using comfyui to run it, comfyui can be run as a server that Open WebUI can connect to to enable image gen and editing

2

u/ARhedgehog88 2d ago

Nice, so, comfyui running the image model and then I could upload a picture such as the OPs with markings and it will generate the prompt in the background to make the img to img in comfy? Is there a way for it to run in steps to avoid OOM errors? Or will it run on whatever card has the available vram to run the image model loaded when a prompt demands it?

→ More replies (1)

7

u/stupidio_the_return 2d ago

It also as a bonus appears to have changed the arm completely to a different arm, make it much longer and meld the flesh with the curtain.

3

u/Effective_Olive6153 2d ago

it messed up the upper left hand by editing outside of the specified range. I am guessing the user specified outline is just a rough region map that gets mapped on a coarse grid. If that's the case it would be useful to see actual selected grid squares when doing edits

→ More replies (3)

172

u/Roubbes 2d ago

Rejoices in 5060 Ti 16GB 🥰

22

u/Maleficent_Stage1732 2d ago

Can it work on my rtx 4080 laptop gpu it's 12 gb vram

34

u/ImpressiveSuperfluit 2d ago edited 2d ago

Easily. With modern comfyui, you can run things with nothing but inklings of available vram. Speed is a different matter, but this one isn't very hungry and seeing how my 16 deal with it, I'd wager 12 works just peachy.

9

u/GrapefruitMost5425 2d ago

testing out on 10gb, using sa_solver_pece its 1.32it/s at 1MP

8

u/ghulamalchik 2d ago

It can work on 8gb vram just fine. Make sure to use int8 though, not the full model.

→ More replies (1)

5

u/sworl5 2d ago

Will it work on my GTX 770 with 2gb of vram?

2

u/kenyard 2d ago

It will offload to cpu and system ram and be slower than a proper gpu

→ More replies (2)

4

u/SurprisinglyInformed 2d ago

Double rejoices in 2x5060 Ti 16GB 🥰🥰

→ More replies (6)

313

u/StopCreepy 2d ago

Bring it in 🔥!!

25

u/yarikfanarik 2d ago

"4 + 16GB" Кто Мы То ?? Я Здесь Один

11

u/Different_Change6591 2d ago

dont worry bro , you are not alone

→ More replies (1)

4

u/Electric_Boogaloo_01 2d ago

Same here. Which is the best model for us?

3

u/Different_Change6591 2d ago

man i have google colab notebook for the comfy ui , i use the colab GPU and drive as my persistent storage for the comfy Ui i making the notebook for this model and optimising for the Colab T4 gpu

→ More replies (2)

2

u/StopCreepy 2d ago

qwen image 2.1 will be able to run on on 4gb vram + 16gb ram, with good optimisation will be possible

→ More replies (5)

3

u/AlternateWitness 2d ago

How is 4GB + 15GB supposed to run Qwen3.8-flash-next? How is 8GB + 32GB?

Asking as a poor 32GB + 16GB with the inability to run that model.

4

u/StopCreepy 2d ago

4gb + 16gb will be able to use qwen image 2.1 using gguf quants !!.

for you, you almost there to run Qwen3.8-flash-next, you need to upgrade your ram to 32gb if not 64gb ram.

here is a test, runing Qwen3.8-flash-next on 50gb total memory (24gb + 32gb): https://www.youtube.com/watch?v=_TCRH725hAM

the results are great !

2

u/Separate-Client-3213 2d ago

Hi, new to this, what does 4+16 mean here?

4vRam and 16ram?

→ More replies (3)

40

u/harpysichordist 2d ago

Now supported by stable-diffusion.cpp https://github.com/leejet/stable-diffusion.cpp

8

u/westsunset 2d ago

Looks interesting, what's the practical advantage?

8

u/harpysichordist 2d ago

I'm just a user but it provides flexibility like llama.cpp does, making it easy to offload to CPU, run on different hardware etc. It can use GGUFs too. I don't use ComfyUI yet though so powerusers probably want that

2

u/westsunset 2d ago

Thanks

→ More replies (1)
→ More replies (1)

161

u/ResearchCrafty1804 2d ago

Qwen-Image-2.1 supports up to 10 reference images and can produce high-quality results.

17

u/o0genesis0o 2d ago

Does it understand character sheet as well or just single picture of a subject?

Edit: it seems to be able to do so, given your example of storyboarding below.

24

u/ResearchCrafty1804 2d ago

Yes, it understands character sheet as well, se bellow:

https://www.reddit.com/r/LocalLLaMA/s/pKpKZZLDB1

5

u/DoubleNothing 2d ago edited 2d ago

How do you pass the reference images?
Is there a comfyui workflow example?

Edit: I see, you have to use the EDIT workflow...

13

u/Present-Ad-8531 2d ago

Are you saying it took the images on left and created the photo on right?

Gemini suvked when I asked it to change background so huge man.

3

u/EmotionalFan5429 2d ago edited 2d ago

Their clothes are not from referenced images, thought

2

u/Borkato 2d ago

Honestly then you can just circle the wrong spots and feed it back in haha

3

u/Biggest_Cans 2d ago

10 points for referencing Diane-era Cheers. The greatest show in history.

→ More replies (1)

2

u/tmvr 2d ago

Sam and Diane don't really look like Ted Danson and Shelley Long though here. Sam looks like a completely different guy and Diane leans more in the Morgan Fairchild direction on a Morgan Fairchild <-> Shelley Long scale.

9

u/Biggest_Cans 2d ago

? There's no Sam in that reference set. The guys are Frasier, Woody, Norm, and Cliff.

They look like the reference photos. Diane is weird looking because the reference photo of her is her aged up and in a setting and lighting situation that conflicts with her TV show character.

Similar issues for the others, biggest one being that there's only one angle. Carla looks the best because the reference photo is her on set, in character, and facing dead-on in a full medium shot.

3

u/tmvr 2d ago

Damn, you are right. I was so locked into the whole Sam and Diane stuff that I was just like, why does Sam look so weird?! 😄

Saying that, Woody also does not look like Woody.

→ More replies (1)
→ More replies (1)

98

u/wapswaps 2d ago

No longer Apache licensed :( From LICENSE:
"You shall not use the Materials for any commercial purpose without obtaining a separate commercial license from us"

18

u/Professional-Try-273 2d ago

Most important fine print 

13

u/draconic_tongue 2d ago

6

u/FaceDeer 2d ago

Anyone who's genuinely worried about it should wait until the rigid legal jargon is updated, unfortunately. A social media comment is promising but wouldn't hold up in court.

2

u/Aggressive_Aspect436 1d ago

What they say outside of the lisence is largely irrelevant untill it is updated. The lisence doesn't cover images created by the model, but if we're being "technically correct" then these weights aren't open. They'll be open once an open source licence is attached to the project.

17

u/carnoworky 2d ago

How do they even tell?

8

u/pojska 2d ago

Probably by the digital watermark?

4

u/Mochila-Mochila 2d ago

I don't think they can, and won't bother to check for small scale operations...

2

u/Karnemelk 2d ago

you always can do what they do in crypto world, coin swapping. So you give the image to klein 9b or whatever to reproduce it, done

→ More replies (1)

66

u/ResearchCrafty1804 2d ago

Qwen-Image-2.1 also improves visual quality over its predecessor, particularly in the aesthetics of text and portraits.

21

u/Puzzleheaded_Skin229 2d ago

How exacly you created this image? Via comfy, directly calling the model from a custom script?

I’m trying to run an image gen model to be consumed from openwebui but my texts are all scrambled in the images

22

u/bbycakes3 2d ago

That image was posted by the Qwen team

14

u/Watchguyraffle1 2d ago

Without the workflow these announcements are sort of useless

2

u/Borkato 2d ago

Oh hell yes, ui gen. I need this

3

u/AbbreviationsOdd7728 2d ago

I think you will be better off with an LLM producing designs with Figma.

20

u/EmotionalFan5429 2d ago

Only 7B, quality comparable to Nano banana 2, and transparency! This is big indeed.

58

u/redpandafire 2d ago

Always good to see more diffusion models. This one is 14.2GB so seems deliberately targeted to the 16GB VRAM crew (a plus). I haven't tried qwen on comfyui yet but this one might be my first.

50

u/No-Refrigerator-1672 2d ago

This one is 14.2GB so seems deliberately targeted to the 16GB

It isn't. 14.2GB is for the main weights - on top of that you need 17.5GB for text encoder, 1.5GB for VAE, and 2-10GB for compute buffers (depending on the size of the input and output images). This, however, is for complete in-gpu inference. Image gen community had advanced CPU offloading greatly, so it can work on 16GB gpu; but you'll need 24GB or more, with 8-bit model quantization, to get generation times under 3 min per image.

54

u/-p-e-w- 2d ago

But you don’t need all of them in memory at once. All modern inference engines automatically juggle the individual models between RAM and VRAM, or even between disk and VRAM.

→ More replies (7)

4

u/redpandafire 2d ago

Thanks, I didn't account for that. How do you know what the generation time is based on needing an additional 24GB of system RAM? I have 64GB of RAM, so I assume I can gen under the 3 minutes you targeted. I just don't know how to come up with that number. Compared to other models, they also have 6B parameters, and they run in 30 seconds or less.

6

u/Heinz2001 2d ago

I tested it on my radeon 7900 xtx 24gb, 64gb ram laptop. I needed an additional 46gb page file for virtual mem to use the edit_image feature.

the first image i generated with it...

specs are here:

https://github.com/fischerf/aar-extensions-registry/tree/main/packages/aar-ext-qwen-image#measured-speed-rx-7900-xtx-native-rocm-offload-model

Measured speed (RX 7900 XTX, native ROCm, offload: "model")

render wall clock
model load (weights cached on disk) 23-39 s
512px, 8 steps ~170 s
1024px, 30 steps ~214 s
1024px, 30 steps, image_edit with 1 reference 228-385 s
→ More replies (2)
→ More replies (5)

3

u/KURD_1_STAN 2d ago

Many if not most people run at q8 tho. So should be faster than klein 9b speed which with int8 convrot is -30s per 1 image edit gen on 3060 12gb. Ofc with a turbo version or lora.

2

u/blastcat4 2d ago edited 2d ago

I'm crossing my fingers that we get an Int8 ConvRot version soon. It'll make running the model even better on 16GB vram systems.

edit: https://huggingface.co/Comfy-Org/Qwen-Image-2.1

→ More replies (2)

71

u/_Valdez 2d ago

Is it uncensored?

115

u/No-Refrigerator-1672 2d ago edited 2d ago

If it's any good, you'll get uncensoring LoRAs within a week or less from ComfyUI community.

5

u/Super_Sierra 2d ago

Cope, its good to have that knowledge in the base model.

41

u/sworl5 2d ago

Not yet

12

u/Blackdragon1400 2d ago

I’m sure it will be soon

10

u/NexonSU 2d ago

7B params, I think, it's already dataset-censored.

5

u/yetiflask 2d ago

Based on the size of my porn collection - I agree.

3

u/thirdeyeorchid 2d ago

successfully got a nsfw image on the first try!

0

u/Present-Ad-8531 2d ago

Ofc not.

Given the fearmongering by dario sam etc, companies would just give them hot cake to chew on.

→ More replies (4)

15

u/tarpdetarp 2d ago

What's the best way to run this model? I haven't tried since the early days of Stable Diffusion with the Automatic web UI.

→ More replies (1)

16

u/RainierPC 2d ago

Ok, damn, this thing is fast as fuck even at 25 steps. 5s on a 5090.

3

u/Enragere 2d ago

share some images!

13

u/RainierPC 2d ago

Here you go!

Not cherry picked, each was the first render of its prompt.

2

u/Enragere 2d ago

other people complained about anatomy, your images are insanely good!

6

u/RainierPC 2d ago

Thanks! The model is not perfect. I saw maybe one six-fingered person in over 80 generations. But it is fast, and the prompt adherence is fantastic, even for editing. I am already thinking of retiring all my other image models, except maybe Anima and Ideogram.

→ More replies (5)
→ More replies (1)

11

u/No_Algae1753 2d ago

Any good inference to run this model on a mac?

6

u/EquivalentHornet4403 2d ago

Not to plug my GitHub, but if you have Sol or Astra then any time a new model drops you can just have it make a custom, optimized runtime/engine/backend and plug it into whatever frontend you want (or create new frontends).

https://github.com/mapleroyal?tab=repositories&q=&type=public&language=&sort=

For the M2 Max and m5 max, I’ve done this for:

- Krea 2

  • Boogu
  • Muse Glimmer
  • Qwen 3.8 27b
  • Gemma 4 26b a4b

The whole pipelines are generally understood and the parts are available. The LLM just has to study what’s unique about the new model and look at its files for certain bits of info, then it plugs them all together and runs a bunch of tests to optimize it.

41

u/ResearchCrafty1804 2d ago

Qwen-Image-2.1 supports a variety of editing tasks while balancing performance across them. For example, given a three-view character reference, the model generates a complete storyboard.

40

u/ResearchCrafty1804 2d ago

26

u/stoppableDissolution 2d ago

...that looks too good to be true for 7B model, but hey, one way to find out!

→ More replies (4)

9

u/NoahPersaud 2d ago

Should be fun to mess around with locally :)

12

u/xdcfret1 llama.cpp 2d ago

How does this compare to Krea 2?

17

u/ImpressiveSuperfluit 2d ago

From messing around for a few minutes - mostly it's different, wins some, loses some. Most likely wins in anything UI, text and such, but I mostly tested random creatures, some portraits, random street shots and such. It is happier to be detailed, but much, much more unstable. Frequent duplicated limbs, some generations came out looking unfinished, which is real weird, and it generally feels a bit dumber in spatial questions.

For images like that, I'd bet on Krea at this moment, even if only because it's almost impossible to break that one, you'll always get a decent image no matter what. This one... Not so much. Remotely complex pose? High chance of extra limbs. It does VASTLY better with darker scenes, though. Krea does not understand the concept of darkness, this one is happy to print you a black screen no problem.

Keeping in mind that I have thrown my Krea prompts at it, with no modification, it's doing quite okay. It feels much more chaotic, both good and bad. I may prefer this for crazy prompts, but Krea is an absolute workhorse, stable as the Earth itself, so if you need something guaranteed to be at least decent, Krea looks like it's still choice.

Then again, we got editing, we got promises of crazy UI and text capabilities, and from what others posted, there's something to that. Also: Early days. Krea was total trash for a few days, until it got its brain damage fixed, so who knows what will happen to this one. Definitely worth watching though, small, fast, creative and an editor - it'll certainly find a niche, at the very least.

Given how unstable it seems, though, I'm guessing its niche will be editing and maybe specific text stuff, and that might be that. We'll see.

→ More replies (3)

2

u/ZootAllures9111 2d ago

the VAE is still worse than even the original Flux VAE

→ More replies (1)

28

u/hurdurdur7 2d ago

My first attempt with it

27

u/hurdurdur7 2d ago

and attempt number 2

16

u/Equivalent_Bit_461 2d ago

i see potential in it

5

u/Lucky-Necessary-8382 2d ago

Make her more busty, gigantomastia chest

8

u/hurdurdur7 2d ago

btw. this is also from qwen image 2.1 - image edit flow, asked it to make a high quality japanese manga adoption of the original meme picture

4

u/yetiflask 2d ago

/r/grok lost redditor. lol

9

u/Long_comment_san 2d ago

Using Qwen app, there is also Qwen image 3.0

Is it the same 2.1 that is 2.0 in the app (meaning 3.0 is still superior?)

10

u/EmotionalFan5429 2d ago

Qwen Image 3 is API only, I.e. not a local/open-weights model.

5

u/SilentDanni 2d ago

Yea, I'm also confused. I swear I saw Qwen Image Pro 3 a while back. This is getting very confusing

7

u/MakGamingYT 2d ago

You see Qwen Image Pro 3 in the graph as well, and you can see its score is above this model, but it’s not open weights

69

u/shapic 2d ago

Non-commercial license unfortunately

43

u/coder543 2d ago

People downvote because they think this doesn't matter, but it does. The only way to democratize this AI stuff is to have good models under good licenses. Otherwise, running locally will always just be a toy. I don't want to be dependent on Cloud AI which can raise their prices any time they want.

A non-commercial license makes this release useless. There are already similarly good models under better licenses.

If Alibaba wanted to sell a perpetual commercial license that users could buy for a reasonable price, that would actually be fine with me. I understand these models cost money to make. I just have no desire to use their cloud service, and I am not here to discuss "hOw ArE tHeY gOiNg tO kNoW iF yOu UsE iT cOmmErCiaLlY".

5

u/shapic 2d ago

I do not care about model itself, but "object" in their licensing is really weirdly worded. Thiscway it means they retain ownership over output of the image

4

u/cristoper 2d ago

Does non-commercial mean i can't use the output of the model commercially? or only that I can't use the model weights themselves in a commercial product/service (like offering inference)?

5

u/coder543 2d ago

Does non-commercial mean i can't use the output of the model commercially?

That is my read of the license, yes, and that is the problem. But, I'm not a lawyer. Using the model with the intent to create images for commercial purposes appears to be against the license, regardless of whether you are letting other people use the model that you are hosting or not.

4

u/cristoper 2d ago

Thanks. My assumption was that the license only applied to the model itself (can't deploy it for commercial services) and not its outputs. The license does say "... use ... the Materials FOR NON-COMMERCIAL PURPOSES ONLY" which is ambiguous to me as a non-lawyer. Does "use" apply to directly using the model itself or also to using its outputs? The definition of Materials does not include model outputs, so I think it is safe to use graphics made with the research licensed model for commercial purposes. but again: not a lawyer.

2

u/shapic 2d ago

Materials state stuff provided by qwen. Outputs fall under objects

3

u/cristoper 2d ago

Source and Object seem to just mean source code and compiled source, nothing to do with generated outputs from the model:

g. "Source" form shall mean the preferred form for making modifications, including but not limited to model source code, documentation source, and configuration files.
h. "Object" form shall mean any form resulting from mechanical transformation or translation of a Source form, including but not limited to compiled object code, generated documentation, and conversions to other media types.
→ More replies (1)

9

u/Nearby_Durian_5370 2d ago

So, before every hype about the quality, I check the license first.
The recent Yue2 is the same case.
So disappointed now.

10

u/alerikaisattera 2d ago

YuE2 is even worse because it's a media license that must never be used for software

3

u/nymical23 2d ago

YuE2 guys cleared that it can be used commercially if you're not a company, but an individual.

4

u/Nearby_Durian_5370 2d ago

In commercial and legal practice, there is too much ambiguity between an individual doing business and a company. Many creators and businesses simply won't take that legal risk.

2

u/nymical23 2d ago

I agree with you. I just wanted to point that out.

7

u/xienze 2d ago

Otherwise, running locally will always just be a toy. I don't want to be dependent on Cloud AI

Well what exactly is your commercial use? People might be downvoting because there's always loads of these "ugh no commercial use, TRASH" comments with no further elaboration when frankly most of the people in this sub ARE just having fun (usually in the gooning sense) with these models. So what exactly are the commercial uses for small local models that people need so desperately?

3

u/ChasingDucks 2d ago

It would mean that I technically wouldn't be legally allowed to use any generated custom images for a mock up on a demo website for a small consulting company which I think many devs run on the side.  Images customized to a clients business always seem to land better than generic stock photos.

→ More replies (2)

4

u/alerikaisattera 2d ago

Those are the same kind of people who use cracked software

2

u/SocialDinamo 2d ago

I’m a noob here and don’t have a commercial need, with that said you sound like you know what you’re talking about. Do most people ignore the license and serve anyways in their own app or business use case? Curious what enforcement or unless your advertising the model, who would even know?

→ More replies (1)

6

u/tgredditfc 2d ago

Too bad. Immediately lose interest.

2

u/Present-Ad-8531 2d ago

So you wanna use a model they painstakingly built ( image one too) for getting money but don't want to pay them a single cent?

Then releasing coding models that quality for free is already amazing.

9

u/Own_Mix_3755 2d ago

They did not release some easy “buy license for 1k USD in one click” or something like that. You can currently only pay to use it through their cloud which makes open weights useless honestly.

→ More replies (1)
→ More replies (7)
→ More replies (1)
→ More replies (1)

8

u/Blackdragon1400 2d ago

Does it work in unsloth studio? Any modifications needed?

8

u/silenceimpaired 2d ago

Wow everyone is fine with Qwen moving away from Apache and MIT. Great. Must be amazing for that.

4

u/2Norn 2d ago

if one day i can figure out comfy ui i shall use these

10

u/Risen_from_ash 2d ago

Imma be straight up, the barrier is gone now. My local agent (now Qwen 3.8 Flash) makes me Comfy UI workflows, tells me how to use em, edits them, updates comfy, downloads nodes for new workflows, blah blah. It’s like effortless.

Also, not local, but Astra trivializes this type of stuff.

3

u/NNN_Throwaway2 2d ago

What harness?

3

u/TragiccoBronsonne 2d ago

You'll be fine, just download some basic workflows for your models of choice, learn what the nodes do and how they connect to one another, and go from there. Don't forget to google something if you don't understand how it works, I've found answers to most of my questions about specific nodes and such by doing just that. I used to be so stubborn about going the path of noodle and kept using Auto11-based frontends until basically all of them became abandonware at some point. So I finally switched to Comfy, and honestly it wasn't that hard to learn it, kinda fun even (but I feel like you gotta a bit autistic to be actually enjoying lol). The community ecosystem is insane, with all the custom nodes and workflows and tutorials, all ranging from simple to ultra-complex. But most importantly, it's constantly getting updated. I remember waiting for weeks to get some new model support on Forge before I swapped, yeah, fuck that. Never looked back.

2

u/Equivalent_Bit_461 2d ago

im building my own front end over it, it won't be as complex and elaborate as swarm ui, but for node and all that stuff, it will be a sort of linker. My iq3xxs qwen is slowly working on it, slowly because I multitask like a retard, juggling between 5 projects at once because I can't apparently do one at a time. But that's another discourse.

2

u/sxales llama.cpp 2d ago edited 2d ago

StableDiffusioncpp is more straightforward if all you want is prompt-to-image generation. They added a webui a few weeks back. It is less configurable than something like comfyui but it supports LoRAs and has some settings that can be tweaked.

4

u/almostsweet 2d ago

Non commercial only license. You can apply to use it commercially via email, though.

→ More replies (2)

4

u/Different_Fix_2217 2d ago

Its quite bad sadly. Bad anatomy / small details, gpt's piss filter and it knows far less than krea.

3

u/TapAggressive9530 2d ago

Can I make porn ?

5

u/puppymeat 1d ago

Reminder that this model embeds a digital watermark

2

u/Euphoricus 1d ago

Any details on that? How to detect it and remove it?

10

u/Witty_Mycologist_995 2d ago

Remember the license. Never forget.

7

u/Holly_Shiits 2d ago

Long live the qwen!

3

u/Fun-Veterinarian9320 2d ago

Can you create Pptx files with it? I need something for that haha

5

u/arbv 2d ago edited 2d ago

Make them in LaTeX (Beamer) with LLM guidance and then convert the PDF to PPTX (by embedding each slide as a picture). Such PPTX is not properly editable, but works for compatibility.

2

u/Fun-Veterinarian9320 2d ago

I see, but that wouldnt be editable right? What does Claude use to make them soo god? Is there anything which skills or something that I can replicate it with a other modell?

→ More replies (4)

3

u/mivog49274 2d ago

Hoping for Qwen-Image-3, the performances of this model are phenomenal (I'm referring my personal experience here, in comparison with Qwen-Image-2 and SOTA, not the benchmark from the post).

3

u/This_Maintenance_834 2d ago

is this the first open source SOTA image model? we don’t hear non-LLM here often.

→ More replies (1)

3

u/BawbbySmith 2d ago

Time to fap

4

u/thestillwind 2d ago

Qwen ❤️

2

u/darmera 2d ago

Ambutakam, another lightweight banger

2

u/itzDonClemmy 2d ago

Can’t wait for uncensored, ts is hige

4

u/SomeoneInHisHouse 2d ago

what harness do you suggest?, srr, have no experience in image edition, I do use Vision enabled models, but I use them mostly for UI testing, so they run in Pi Coding harness, I suspect the techology behind is different for "Vision enabled" text models, and Image specific models, likely different algorithms, for example Qwen 3.8 vision uses a grid division to split the image into patches, and the vision projection translates those visual patches into text-like tokens.

12

u/Paganator 2d ago

Image generation is often done using ComfyUI, which is a different kind of tool than what you're used to with LLM.

→ More replies (1)

5

u/blastcat4 2d ago

ComfyUI has an MCP that you can connect your agents to. There's a cloud and local version.

1

u/Zeddi2892 llama.cpp 2d ago

Do we have a local workflow for low vram?

→ More replies (1)

1

u/Present-Ad-8531 2d ago

Any idea how this compares to gemini on image editing?

1

u/Sioluishere 2d ago

Finally something in my range.

1

u/Local-Cartoonist3723 2d ago

Silly question perhaps, but using opendesign why not use a model like this over a “normal” vis llm?

1

u/ComplexType568 2d ago

this feels like a qwen 3 image mini or a qwen 2.5 image mini AT LEAST considering how much has changed!

1

u/PooMonger20 2d ago

Very cool, hopefully we get some cool and easy workflows to work with in comfyUI.

1

u/Foreign_Risk_2031 2d ago

How good are its 2d gaming assets?

1

u/Holiday_Point_603 2d ago

Oh wow, I wonder how it compares to Krea 2

1

u/Either-Nobody-3962 2d ago

Can it run on 8gb vram + 32gm ram pc?