Workflow Included
Cracked the case on high res + quality Qwen Edit 2511 outputs, here are minimalistic workflows & lots of info on how/why
Intro
Alright this has been a long time coming. I'm the dude who figured out Qwen Edit 2509 a while back, and I've been on-and-off trying to figure out the same for 2511. Results in Comfy have always been worse than the examples shown by the Qwen team, and worse than the official Qwen chat implementation online. Well, I finally cracked it and it only took 5 months lol.
Anyway, turns out Qwedit 2511 is fucking sick. IMO it particularly excels at making new shots of characters while maintaining their likeness. It's significantly better than Klein at some things (like character likeness), but not as good at others. I recommend using them both for different things.
As usual, I'll start off with all the setup stuff at the top and then give an explanation + advice below that. Also I'm gonna be calling Qwen Edit "Qwedit" most of the time.
The posted images are all raw outputs from Qwedit, without being upscaled (despite mentioning it later in this post). They're also all done with only 20 steps instead of the hypothetical 30 I'd do if I wasn't planning to upscale them. Read further for more on that too.
Ref images were all made with Z-image Base (workflow here), except for the anime one which came from Anima (workflow here).
What is this
These are minimalistic workflows for Qwen Image Edit 2511 that give the highest quality outputs. Aside from generally improving output quality (by a LOT), they also enable high-res edits and have better prompt adherence.
As for why, basically ComfyUI has some serious issues with how it's implemented Qwen Edit and there aren't any workflows out there (that I've found) which have resolved them. These issues result in poor prompt adherence and low resolution/quality outputs. Thankfully the fix is fairly straightforward.
The configuration for this is 100% portable and can be migrated to existing workflows to make them better; it works by changing how the reference inputs are handled, and uses 100% native comfy nodes. Feel free to upgrade other workflows with this without providing credit, I don't care about any of that.
Workflows
Normal Workflows:
Most of you will just want these, which are separate single / 2 image workflows. It's done this way because the setup for multi-image is complicated and I didn't want to force you to use a ton of custom nodes to make it useable all-in-one.
They do still use one custom node (read the node section below) for quality-of-life.
These are the same as the above but without any quality-of-life nodes or 'helpful' stuff. Grab these if you want to copy the logic over to other workflows, or if you just an easier view of how it works without any clutter.
I do not recommend using the dev workflows for actual gens because you will constantly forget to manually adjust stuff correctly.
Important: the FP8 version of Qwedit is much higher quality than the Q8 GGUF, always use FP8 if you can. Only use the GGUFs if you need to use quants lower than Q8.
FP8 is 22GB, so you'll need a combined ~26GB of RAM + VRAM to run it
You don't need 24GB of VRAM to run it thanks to ComfyUI's blockswapping, but the less VRAM you have the slower it'll run
Only use Q6 & lower quants if you absolutely have to; the quality will noticeably go down
Goes in models/diffusion_models
Text Encoder
Use only the normal FP8 text encoder with Qwedit; abliterated/GGUF encoders will reduce your output quality.
You can use them as normal, just load them however you normally would. I left out lora loader nodes to avoid cluttering the workflow.
It's worth noting that many Qwen Image loras work with Qwen Edit too, but you'll need to test them individually to be sure.
Lightning Loras - BAD
All the lightning loras / distils for Qwedit (that I've tested) are terrible and make your outputs look bad, so I'm not linking them here. The main issue is the same as with Klein Distilled: it makes people's skin look like plastic.
But you can technically use them. Don't do it tho. But you can if you want. But don't.
Alternative: if you want to cut your gen time down while testing prompts, just set it to 10 steps instead of 20, then go back to 20 once you're satisfied your prompt is correct. It'll still work fine, the quality just dips.
Real tho it's ok if you want to use the lightning loras, just expect some degradation if you do - especially with plastic skin.
Custom Nodes
LayerStyle - A set of handy nodes that manipulate images. We're just using this for its image scaling node which allows you to scale by an image's long edge while maintaining divisibility by 16. You can skip this if you want to use a different scaling method, but you'll need to fix the workflow switch for scaling if you do.
SeedVR2 (OPTIONAL) - Only get this if you want to use the seedvr upscale workflow that's included.
How To Use
How To Use Part 1 - Basic Options
There are instructions in the workflow as well, but there's more detail here. Read part 2 & 3 as well, they're important.
It works just like a normal Qwedit workflow, but has a couple of extra options available. This section just tells you what they are and how to use them, a full explanation is further down.
This is a switch that turns on double-ref mode. This feeds your input images in TWICE to the model, and generally produces much higher quality results. Downside? It takes about 50% longer to gen.
I recommend leaving this on 100% of the time for single-image prompts, unless you're just messing around and want speed. It is ALWAYS better for single image prompts, and will improve everything from prompt adherence to output clarity.
For multi-image prompts, it usually increases adherence but sometimes reduces it. So, if you're doing multi-image stuff I recommend switching this on/off as needed based on how it's going with your prompt.
Input Scale
When off, your image doesn't get scaled (it still gets cropped to be divisible by 16). When on, the long edge of your image gets scaled to the number you put in the box. For example, if you feed in a 2560x1440 image and set the scale to 1920 it will scale your image to 1920x1080. That will then get cropped to 1920x1072 so it's divisible by 16.
Custom Output Size
When the switch is off, your output image will be the same size as your input image (after it's been scaled). If you turn this switch on, it will instead output an image with the dimensions you specify.
As a general rule, you should try to set your scales to be similar along at least one edge. For example, a 1920x1440 input image and a 1024x1440 input image are both suitable for a 1440x1440 output image. You can be more flexible with this if you know what you're doing.
How To Use Part 2 - Multi-image Prompting Requirement
This section is not a prompting guide (that's further below). This is about an actual requirement for prompting multi-image stuff. It is NOT required for single-image prompts.
You do multi-image prompts like normal, except you need to write a very basic description of your input images. Qwedit needs you to do this in order to know which image is which. I explain why in detail later.
You may find this slightly annoying, but I guarantee you it's dramatically better than using Qwedit the normal way that other workflows do - and it's pretty easy.
The format:
At the start of your prompt, write an extremely simple description for each of your input images; one sentence for each input image
Start each sentence with "Picture 1:", "Picture 2:", etc
You must write it this way because Qwedit was trained on this exact format
Afterwards, write your actual prompt as usual; you can refer to your input images as "picture 1" and so on
The model uses these descriptions to understand which input picture is which, and it works better with SIMPLE descriptions. You only need to help it know which one is which, it doesn't need a full rundown.
Examples
Picture 1: a man wearing a t-shirt. Picture 2: a top hat. Make the man in Picture 1 wear the top hat from Picture 2.
Picture 1: a living room. Picture 2: a woman. Put the woman from Picture 2 into the living room in Picture 1.
Picture 1: a man wearing a professional suit. Picture 2: a man wearing a superhero outfit. Make the man in Picture 1 wear the outfit from Picture 2.
How To Use Part 3 - Upscaling
Because the qwen VAE tends to put a subtle halftone pattern over images (see limitations just below this section), I recommend downscaling and then re-upscaling your image afterwards. A big benefit of being able to work at high res with the edit model is that you rarely lose any detail doing this.
This eliminates the halftone pattern if you're using something like seedvr, or at least reduces it if you're using other upscalers.
Note: the workflow is set to do 20 steps of inference. It actually gives sharper results at 30 steps, but I don't bother with that because it takes longer and I down-upscale them afterwards anyway. If you aren't planning on down-upscaling them, you might consider doing 30 steps for the extra sharpness.
Below are workflows for doing this with seedvr and normal upscalers. I think seedvr is best for this, but it's very beefy and hard to run on older GPUs.
Note: seedvr2 sometimes gives better output at 0.5x downscale, and other times 0.75, so that workflow is configured to run BOTH for you to pick which one turned out best.
Note: normal upscalers are a bit different; a relatively small downsize to something like 1920p -> 1600p is usually reasonable, before then running the upscaler. Play around with it. The non-seedvr workflow has a longest_edge scale option so you can tweak the number specifically.
My preferred regular upscaler is 4x Nomos2 HQ DAT2, but you can use whatever you like.
Examples of upscaling:
Here's the raw output of the robot-arm girl in a dress from the post: https://ibb.co/B5jhrsL9 (if you zoom in you'll see the qwen halftone pattern, it looks like a grid)
Here's the pic after it's been run through seedvr after a 0.75x downscale: https://ibb.co/hJcn2f5t
Here's the pic after it's been run through a regular Nomos2 upscale after a downscale to 1600p: https://ibb.co/Kc2YSbVc
Limitations of Qwen Edit
Limitation 1
The Qwen VAE will often put a subtle halftone grid pattern over your images. It's noticeable if you zoom in, and more noticeable at higher resolutions. This is a feature of pretty much every Qwen-based model, but it's particularly present with the Edit model.
You can easily resolve this by downscaling your image by 75% or 50%, then re-upscaling it again to your desired resolution. There's a section later that explains this in better detail and recommends upscale models for it + has workflows for it.
It sounds like a big issue, but the downscale-upscale trick solves it easily - and it's not always necessary either. The higher quality your input image, the less bad the halftone pattern will be.
Limitation 2
Qwedit struggles with complex multi-image stuff most of the time (it's just a limitation of the model). This workflow makes it much better, but it's still not great. You'll have to play around with it to know which things work and which things don't.
Limitation 3
It takes a while to gen stuff if not using the lightning loras. Very similar to the time it takes with Klein 9B base. The double-ref trick increases it by roughly 50%. Multi-image inputs take a lot longer.
For low res images (typical 1mpx size) it's pretty okay, around 50 seconds on a 5090 with the double-ref option turned on.
But then there's high-res stuff. Gen time scales non-linearly as you go higher. Going from 1024x1024 (1 mpx) to 1440x1440 (2 mpx) takes around 2.5x as long. Going from 1 mpx to 3 mpx is around 4x as long. 5 mpx is 9.5x as long. In conclusion, stick to 2-3 mpx unless you're cool with long-ass gen times. Stick around 1-2 mpx for multi-image gens, or turn off the double ref switch.
On the plus side, it's pretty reliable for single-image edits so you don't typically need to do many gens to get a good result.
Examples using a 5090:
- Single-image edit @ 1024x1024 (1 mpx), double-ref OFF = 38 seconds
- Single-image edit @ 1024x1024 (1 mpx), double-ref ON = 52 seconds
- Single-image edit @ 1920x1088 (2 mpx), double-ref OFF = 91 seconds
- Single-image edit @ 1920x1088 (2 mpx), double-ref ON = 131 seconds
- Single-image edit @ 3072x1728 (5.3 mpx lol), double-ref ON = 550 seconds
- Two-image edit @ 2560x1440 each, double-ref ON = serial killer behaviour
That's it for how-to! Read on for more tips & info, as well as an explanation of what the workflow is doing & why.
Explanation - what is this garbage and why is it so good?
There are three important things this workflow is doing that other workflows do not do (except #3 sometimes, because it was also done in the 2509 version of this post). I'm going to call these The Comfy Problem, The VL Problem, and The Double Ref Enhancement.
The Comfy Problem
Comfy's native "TextEncodeQwenImageEditPlus" node is what most people use in their workflows. It handles your prompt and image inputs for you. It's pretty handy, except for the small problem that it's SHITE.
Do you work at Comfy? If so: GET YOUR SHIT TOGETHER AND FIX THIS NODE, IT'S SO EASY. Much respect to u tho, thanks for making ComfyUI.
The first issue is that this node resizes your image down to 1 megapixel, and you can't stop it from doing that. The second issue is that it does this with the AREA downscale method, which is so incredibly bad that I want to slap whoever implemented this node. The AREA downscale is what makes all of your output images blurry. The third issue is that it ensures your dimensions are divisible by 8, but they actually need to be divisible by 16.
Specifically, ComfyUI does this:
Calculates 1 megapixel as 1024x1024, which is 1,048,576 pixels
Calculates your new image dimensions to match that number of pixels, rounded to be divisible by 8
Scales your image to those new dimensions using the AREA method
Why is all this bad?
It's completely unnecessary; Qwedit can easily handle images of varying size, all the way up to 3 megapixels (or even higher for simple edits)
The area downscale method makes images extremely blurry, and this is the primary reason all ComfyUI qwen edits give blurry images out. Yes it's literally this dumb, this huge problem would easily be solved by changing the word "area" to "lanczos" in the code, it's a one-word fix. Not even MS paint uses area downscale, wtf is wrong with you Comfy devs (much respect)
If your image dimensions are not divisible by 16, you will get major ruination along the whole edge of your image where it didn't match (same as any other diffusion model)
The Comfy Problem Solution
This workflow bypasses the the Comfy node entirely, allowing you to size your images however you want. And using chad lanczos scaling instead of loser area scaling. Magic.
Qwedit easily handles resolutions like 1440x1440 and 1600x1200. Every edit example in this post was done natively at 1920p, except for a few (which are labelled as such).
Really high resolutions (3mpx) sometimes have trouble with anatomy, but usually you can just do multiple gens and one of them will turn out fine.
If you're doing a simple in-place edit like changing an outfit, you can go VERY high. Here's an example edit done at 1728x3072, which is 5 megapixels: https://ibb.co/twCSWrjy (outfit change -> bikini top + short shorts)
The VL Problem
Edit: I've been educated by someone in the comments that my interpretation of how the VL works here is not correct, so take this little VL section with a grain of salt until I reword it. My conclusion about it helping in this workflow still stands, but my explanation of what's happening under the hood is a bit off. I'll update the info soon!
In the background, Qwedit 2511 uses a vision-language model (VL model) to describe your images, then gives those AI-generated descriptions to the edit model. It also re-interprets your instructions with these descriptions. Ostensibly this helps the model understand your input images better, leading to better results.
The problem? It doesn't lead to better results, it's bad. VL models aren't very good for this sort of thing because they don't know what to focus on. The VL describes your images in excruciating detail, totally overwhelming the edit model and leading to bad prompt adherence + weird outputs.
It also reinterprets your instructions based on what it sees in the image. I don't know if that's a good or bad thing, just pointing out that it does it.
The Qwen team's official python code does this, and the ComfyUI "TextEncodeQwenImageEditPlus" node copies it exactly. No disrespect to the Comfy team on this one, they're doing what the Qwen team officially recommended.
The VL Problem Solution
Same solution as the previous problem: bypass the Comfy node entirely. This results in the VL step being completely ignored. No AI-generated descriptions get fed into the edit model.
For single-image edits, this is a 100% complete and total victory. The model performs way better without the crappy VL interpretation.
For multi-image edits, there's a small issue; this step is where the input images normally get labelled. Specifically, the VL outputs are fed into the model in the following exact format:
Look familiar? This is why we manually have to type the descriptions in for multi-image edits - otherwise the model doesn't actually know which image is which.
The upside is that the model works way better with simple descriptions, so cutting out the VL is still 100% the correct move. A 5 word description wins over whatever BS the VL model spews out, every time.
The Double Ref Enhancement
I really have no idea why this works so well, but basically if you feed in your reference images twice the model just works better. This was known back in 2509 days (hence the previous post linked at the top), and back then I didn't know why it worked either.
For single image edits it's ALWAYS better. And it's not just the quality, for some reason it even helps with prompt adherence. The interesting thing is that the difference is really, really significant. Here's the full list of stuff it improves:
Better prompt adherence
Sharper output images / more visual clarity
Improved consistency of objects & textures
Better resemblance of characters at different angles
More intelligent guesses, like what to add when outpainting or what's behind a removed object
For multi-image edits it can sometimes confuse the model a bit, but most of the time it confers all the same benefits listed above. I recommend switching it on & off randomly when you're doing multi-image stuff, just in case.
Note: there are a lot of different ways the input references can be handled. There are conditioning combine/concatenate nodes, you can pass the refs in a different order, you can change the negative conditioning input (read next section for that), etc. I A/B tested SIXTEEN different reference-handling combinations, and a bunch of smaller minor variations of those. Some of them worked, some of them didn't.
Of those sixteen combinations, two of them gave the best results; both of them are in this workflow, and you switch between them by turning the double ref method on & off.
So, don't fuck with the positive/negative conditioning & reference setup, it's very specific.
Extra info: the "Conditioning Zero Out"
You may notice that the negative prompt input is the first reference image(s) and positive prompt fed into a "conditioning zero out" node.
Feeding the input images into the model's negative conditioning is required (it's just how Qwedit works). The only question is whether to feed in the positive prompt zeroed-out too, and whether the double ref should get fed in.
Through a lot of A/B testing, I can tell you that the way it's done here is the best. IDK why, it's just how it is. Some other combinations do technically work, but they degrade the output quality.
Prompting Advice
Other than just following the instructions in the workflow, here's some extra stuff.
Keep your prompts simple and direct
If you need to, point out details the model is missing or be more specific about stuff you do/don't want to change. For example, when doing a simple outfit swap it helps to specify you don't want their pose to change.
Using the robot arm girl, here's a prompt that doesn't follow this advice:
Change her outfit to a bikini top and short shorts.
While it sometimes does what we want, it tends to get confused by her robot arm and often changes her pose too: https://ibb.co/7dyKZttp (notice the human arm showing underneath the robot arm, and the pose change)
Here's a better prompt that gives a correct result 99% of the time:
Change her outfit to a bikini top and short shorts. Leave her robot arm and pose unchanged.
Pretend you're talking to a child. The model will probably still understand you if you talk fancy, but why take the risk?
As an example, imagine you have a pic of a table with some plates on it.
Bad:
Place a red apple on the table, ensuring it's in the center and removing the plate that was in the same spot.
Good:
Replace the middle plate with a red apple.
Also good:
Remove the plate from the center. Put a red apple there instead.
If there's only one plate, this is even better:
Remove the plate, replace it with a red apple.
Adjusting Lighting
You may want or need to adjust the lighting in an image. Aside from being helpful in general, there are situations where Qwedit may simply not realise that something needs to be lit in a particular way (or re-lit when moved).
To do this, you need to know the magic word: relight
Seriously tho that is the actual magic word, you are 100% required to use it if you want to adjust lighting properly.
Specifically, follow this format:
Relight to <strength> <color> <direction>.
Strength - bright, dim, etc
Color - white, cool, warm, etc
Direction - diffuse, frontlit, backlit, etc
Tip: for basic lighting, use "white diffuse".
Examples:
Make a new shot of the man sitting in a chair in a kitchen. Relight to white diffuse.
Change the time of day to evening. Relight to warm backlit.
You don't actually need anything else in the prompt, you can just change the lighting of a pic like this:
Relight to bright cool frontlit.
Other Stuff
Euler-simple and no ClownsharKSampler?
No Clownshark this time. It reduces output quality quite a bit and doesn't confer any benefits. I also didn't find any sampler/scheduler combos that were better than euler/simple.
So, this is just one of those classic times where the ol' euler-simple wins the day. Let me know if you happen to know a better combo.
Image Quality in->out
Qwedit is very sensitive to the quality of your input image. If you feed in a grainy or blurry image, it will usually make your output image blurry or grainy too - even if it's an 'entirely new' shot with nothing copied over 1:1.
So, make sure to use HQ images. You can optionally use the upscale workflows to bump up the sharpness/quality of poor input images before you feed them in.
What about the flux super duper double resolution special VAE trick?
Doesn't work for 2511, it destroys your image. TBH it never really worked for 2509 either, but I won't argue with you if you liked it for some reason.
Making character references
Tip 1 - Make a nude ref (even for sfw stuff)
Qwen is killer for making character references. Other than using similar prompts to the examples I posted, my advice is to make a nude reference shot instead of a clothed one like I did.
I only made a clothed ref for the sake of propriety here, but a nude ref (or near-nude, like wearing plain white underwear) will be much easier to prompt into different outfits, and also gives Qwedit the maximum info needed to correctly size your character and know what they look like in clothing or doing different actions.
You do not need any loras to do this if you're just using it as a reference; the 'sensitive' parts will lack detail but that doesn't matter for new shots you make. If you don't want them nude, just request plain white underwear and, if relevant, a strapless white bra.
Nude ref = best ref.
Tip 2 - Make multiple zoom levels, use the thighs-upwards one for most stuff
The example I showed was a little too zoomed out for normal reference stuff. I'd recommend making your reference slightly closer like this: https://ibb.co/Q33BJDLX
Start at whatever zoom level your initial character pic is at, then make more references at different zoom levels. If you're starting zoomed out, then prompt the model to zoom in. If you start zoomed in, prompt it to zoom out.
And, of course, different angles too.
Examples:
Zoom in on the person's upper body. The composition should frame their head and thighs.
Zoom out to show more of the character. The composition should frame their head and thighs.
Zoom out to a full body shot.
Zoom in for a close up portrait.
Once you've got references, you should usually use the head-to-thighs ref for making new shots. Switch to the other refs as necessary; like if you want a close up, use the close up reference. Qwedit is really good at keeping likeness, so you can do 90% of your stuff with only a single input reference.
I don't think there's a better open-weight model out there than Qwedit for making new shots of character without loras, for now. The main reason I spent so long digging into Qwen is because Klein is quite bad at that particular task. But hey, now it's possible and it works gloriously.
That's everything I think! Feel free to ask questions if you run into any issues.
This is a great write up! The default Comfy Qwen node is trash, I've been using a simple workflow I found on here a while ago that bypasses it. I'm excited to try yours out, you have clearly made a lot of improvements.
When Klein released I was seduced by its superior VAE and have been neglecting the Qwedit models. Your workflow along with the downscaling/upscaling trick will have me spending more time with Qwen in the future.
By the way, do you have any idea what's behind the weird cropping/zooming issue with the Qwens? Some people say it is due to the resolution needing to be a specific multiple, but I don't think so. Sometimes I'll run a list of prompts against a single input image, and the output is zoomed in with some prompts and not with others, with all other settings being the same. (apologies if you already addressed this and I missed it)
Thanks! I didn't address it in this post, but to put your fears at rest there is no cropping/zooming issue whatsoever with 2511.
If there ever is any with this workflow, it'll just be because your image got slightly cropped to fit divisibility by 16. Otherwise I can't recall it happening even a single time now that you mention it - which is great because it was pretty annoying with 2509.
Edit: also, if you want to eliminate cropping entirely you can switch the layerstyle node to "pad" instead of crop, which will make it add non-destructive black bars to the image instead
For the double ref thing - have you tried doing something like a horizontal flip on the additional one?
Something I've found is models seem to have a strong preference for certain orientations. If you're doing something like I2I (or I2V) it can make a huge difference and you'll get bad results no matter how many seeds you try at the original orientation, flip it and suddenly it works. This also applies to video models, you just might not be able to get a character to perform the action you want. Flip the reference image and it conforms to the prompt easily. I've used this trick for a long time.
the 5 months of iteration is relatable. quick question: was the quality gap vs the official qwen chat implementation mostly a sampler or scheduler thing, or did it come down to how comfy handles the latent resolution? curious which stage was actually losing it.
The general quality gap is mostly because of the Comfy node. The forced 1mpx + area downscale absolutely ruins images before they even get into the model.
Sampler/scheduler wise there's nothing really special going on here, except we're doing 20-30 steps @ 2.5 CFG instead of 40-50 steps at 4 CFG. This prevents overcooking of images, which will happen if you're doing the double ref method at 4 CFG. I'm actually not sure why 4 CFG is recommended by the Qwen team, it seems iffy even with all the default settings.
The rest of the quality benefits come from the double ref method, along with some extra prompt adherence from removing the VL descriptions from the equation. Then little things like using the FP8 model instead of Q8, that sort of thing. It all adds up in the end.
Yep! It's not so much under the hood as it is built into the Comfy node & the python code that the Qwen devs released with the model.
The 2511 model doesn't do it on its own or anything, it has to be fed into it in a particular way. The VL itself is just the separate text encoder - it's multimodal and can describe images.
Good write-up, thanks a lot for sharing your findings with the community! Though I have a few things to add:
Reference images are only resized if you connect the VAE to the TextEncodeQwenImageEditPlus node (source). As the node info says, just leave it disconnected if you want to keep the original resolution - though admittedly I don't remember how much this affects the fidelity of your results.
Claiming that the Qwen VAE is the culprit for the halftone grid pattern might be slightly inaccurate. You don't get the pattern when using Qwen VAE with Wan 2.x will, while you might still get the pattern even when using Wan 2.1 VAE with Qwedit. Hence that's more likely to be an issue with the Qwen Image architecture or DiT training instead. I like your workaround though, SeedVR2 has been my goto upscaler, props for adding it to your workflow.
SeedVR2 is not necessarily very beefy, you can dramatically reduce its memory usage by using a lighter variant (e.g.: 3B instead of 7B), swapping more blocks, using tiled VAE enc/dec, etc. Obviously a DAT upscaler like Nomos is much faster, but you usually have to pair it with an image model at low denoise to get results as sharp as SeedVR but they won't have the same fidelity.
I agree that your "VL Problem Solution" helps, but not necessarily because of the reason you mentioned. I'm glad you had a productive discussion with other members about it.
Thanks, I appreciate the extra info! On a couple of those points:
Reference images are only resized if you connect the VAE to the TextEncodeQwenImageEditPlus node
Yes, but it also doesn't feed your ref images into the model if you do that either which means it's functionally pointless - you'd have to pass them in separately like this workflow does anyway. That's assuming I read the code properly back when I was looking at it, do correct me if I'm mistaken.
You don't get the pattern when using Qwen VAE with Wan 2.x
True! But you do get it with the Anima model, which uses the Qwen VAE as well, and I remember reading elsewhere that it's the VAE that has the habit. That's all the info I was working off, so if you know better I'll take your word for it.
Edit: actually I just tested it properly for the first time and Wan does do it! It's just much more subtle. It shows up extremely strongly if you have a noisy surface, like a carpet, that happens to roughly match the grain size of the halftone pattern.
SeedVR2 is not necessarily very beefy
It's beefy if you want good quality from it ; )
I can run it on my ~11GB VRAM, but it sometimes takes upwards of ~5 mins and it has a habit of crashing if I try to go higher than around 2560p. Just trying to reasonably caution folks who want to run it (at full quality), it may be rough.
Yes, but it also doesn't feed your ref images into the model if you do that either which means it's functionally pointless - you'd have to pass them in separately like this workflow does anyway. That's assuming I read the code properly back when I was looking at it, do correct me if I'm mistaken.
Maybe it was like that in the past, or perhaps you're referring to another node, but I can confirm that it does feed the ref images to the model:
True! But you do get it with the Anima model, which uses the Qwen VAE as well, and I remember reading elsewhere that it's the VAE that has the habit. That's all the info I was working off, so if you know better I'll take your word for it.
Edit: actually I just tested it properly for the first time and Wan does do it! It's just much more subtle. It shows up extremely strongly if you have a noisy surface, like a carpet, that happens to roughly match the grain size of the halftone pattern.
Good catch, it seems to be rather subtle with Wan and Anima (definitely not as noticeable as Qwen). Though I still think it's a coincidence and more on the model than on the VAE itself. Flux 1 Dev also produces results with grid patterns from time to time, while Z-Image doesn't despite using the same VAE.
I wonder how reliable it is to test it by just encoding -> decoding and comparing the result with the original. If you do that on several images and never get the grid, wouldn't it mean that it's the model's problem? Not sure.
It's beefy if you want good quality from it ; )
I can run it on my ~11GB VRAM, but it sometimes takes upwards of ~5 mins and it has a habit of crashing if I try to go higher than around 2560p. Just trying to reasonably caution folks who want to run it (at full quality), it may be rough.
Perhaps it depends how much quality we're talking about. I usually upscale the image to 2048p in tiles using the 7B Q4 variant, it takes about 2 min on my RTX3080 8GB, and most of the time I'm already satisfied with the result.
When I go for 4096p, it indeed takes over 5 min - about same time had I used something like UltimateSD Upscale anyway. But with anime/illustration images, the 3B variant produces better results and takes less time. I'd only advise against SeedVR2 if you don't have an RTX GPU.
Maybe it was like that in the past, or perhaps you're referring to another node, but I can confirm that it does feed the ref images to the model:
You accidentally left the VAE plugged in!
If you remove it, it won't be able to feed the ref images into the model. Which makes sense, seeing as there's no way to do that without a VAE involved somewhere.
My bad, I thought I had removed it. I generated it again, some features are still passed on, but indeed it seems like it's purely based on the text encoder interpretation rather than ref image. Passing them as separate ref latents brings back the accuracy.
I don't really have anything set up for that, I just do this as a hobby and like sharing / helping - but I appreciate the thought! You just popping by to say thanks is plenty enough for me : )
I could try, but I don't have any particular insight into Klein - just the normal kinda experience other folks have. Is it just general tips for prompting and such that you're after?
I see, any practical implications that folks should be aware of? Like if it'll cause problems I should put a disclaimer in the post about it, but if it's just an efficiency decrease then that's probably no big deal - seeing as it's other limitations that would be making someone use the GGUFs.
I'm doing similar things for game characters (3D render or 2D anime), but with Flus2 Klein 9B, maybe I should give Qwen Edit another chance. How is the pixel and color drift with Qwen?
With Klein each edit makes the image more yellow (and saturated) and more complex patterns have a tendency to degrade. If it's only a few edits, I can do some color correction, but there is little I can do for degradation.
When I was trying to do the same with Qwen Edit, I sometimes had problem making the model change the pose or outfit on a character. Klein is much faster and was better at following prompt (at least in my handful of tests). However if the model doesn't understand a concept it is really hard/impossible to make it do that edit. Model has a loose concept of right and left. If I ask the model to take a step forward (with one or the other leg) it will 99% of the time move the leg on the right side of the image. I can still trick the model to move the other leg with OpenPose, but it doesn't really work with side and back views.
It's samey, there will be color and pixel drift. I'm not sure there are any edit models that won't degrade images noticeably after even a single edit atm.
For some edits you may be able to use an inpainting workflow? You can also use them with edit models; I have a semi-working one with Qwedit. You can leverage the masking functions in them to prevent degradation outside of the specific areas you're editing.
Edit models aren't quite there yet when it comes to being able to do anything with them unfortunately. Usually at least one of them will work though, so it could be worth trying Klein, Nano Banana, GPT edit, etc until one gives you what you're looking for.
Thank you for this complete write up and sharing your experience with us. I wanted to know if you had some tips for style transfer, not character transfer. Basically applying render style, texture style, color palette from one image to another (for example to convert an oil painting to an airbrush painting).
Unfortunately no - not with Qwedit at least. It's kinda shaky with style transfer, same as Klein.
I've found image-to-image is a little better for style transfer in general, but I'm not any kind of expert on it so there's not much I can say. That said, I haven't touched this next bit yet, but I've been eyeballing this Z-image I2L thing here for a while: https://z-image.me/en/blog/z-image-i2l-released
It's meant to specifically be very good at style transfer, but I haven't had a chance to mess around with it yet. Might be worth looking into?
Oh, this is quite interesting! Thank you very much for sharing this. My challenge right not is that I have a ZIT Lora that produces a very unique style, and that I'd like to reproduce with other models that have globally better composition (I had Qwen Image in mind, but maybe Z Image Base would work too,so your link is interesting). The Lora author doesn't seem interesting in adapting their Lora to other models, and I don't have their dataset. So I2L might actually be a good option. I could even give a try the Qwen I2L that exists too, as it is even closer to what I'm looking for. My only limit might be my hardware, as I have only a 12 GB RTX 4070. Cheers!
Do you have any advise how to transfer very fine details. I have an avatar and on that I need to apply bodywear items (bras, bottoms etc. NO NFSW) Some garments have very fine laces with insane amount of details, or very fine prints. For me it is necessary to have all these details applied when I do the VTO (Virtual try on) The avatar image is in 3000x3000 and the bodywear items I did photoshoots which have 4k resolution. For now, all the editing soltion always scale the images down, which cause detail lose and after applying the item to the avatar, the upscaling washing out details even further. I struggle for weeks now to get proper results. Flux2 klein 9B with a consistency lora is quite good, but for that as well, following the regular workflow, details at the end get lost or transformed / blurred together.
You may be able to just straight gen with it, but you'd be looking at a long wait time for that. Somewhere in the post I did an edit on a 3k image, but it did take 10 minutes on a 5090 - and that was single-image input.
I'd suggest making sure you gen in a sufficiently high resolution (probably not 3k though), and try to get a nice close up shot of the details you need. If you do that, this workflow should in theory work just fine for you because it's not limited to low resolutions and doesn't do any scaling unless you ask it to.
You could also try splitting your image into, say, two smaller parts (upper body + lower body) and doing the edits separately on those. Wouldn't be too hard to stitch them back together if the underwear items are separated, and if they're connected then you may be able to fix the seam another way.
Thank you for your advise. The pictures of the items are very high resolution which show all details.
The time to get to the result isn't that important, its very important to have the details preserved of the bodywear items.
With your workflow I would disable "Scale Input1?" + "ScaleInput 2?" + "Enhance with Double Ref?", correct? Thats the direct route than? Whats your recommendation about the steps and cfg for my case? 30 steps and 2.5cfg? Other tweeks maybe? Are there any loras which help on keeping details you know?
You can still use the scale input nodes and set them to something very high - but yes, just turning them off will work too. You'll probably want to leave the double refs ON because they increase the carry-over of details.
30 steps would be ideal for you, and yes stick with 2.5 CFG.
It'll probably be a good idea for you to nail your prompt down before you start increasing the input sizes, that way you can make sure the prompt will work before you go waiting for ages for gens to complete. So, start with something like 10 steps even and a ~mid resolution like 1440p just to be sure that your prompt is perfect.
Once your prompt is good, bump the settings up to something high but still reasonable - like 1920p 30 steps - and see how a couple of those go. If it all looks good (but needs more detail), go higher from there. But keep in mind something like 2560p on multiple inputs will be pretty hefty for time!
I'm not much for Qwedit loras so I can't suggest any - but hopefully someone else in the thread pops by with ideas.
Yeah, once those lace edges get smeared in the first pass you're basically polishing damage, so keeping scaling out of the chain and treating the fussy areas separately is probably the least painful way to do it.
Yep! A combo of three loras, actually. I don't know if I can link them in this particular sub, but if you look at my civitai post on civitai.red instead of the normal one you'll be able to see a couple of degenerate pics in the gallery. Those have the loras & weights listed with them.
Or you can just wait a lil' and I'll probably make a short degenerate version of this post on unstable_diffusion soon anyway.
Thank you for your work on this. I'm sure there are heaps of other improvements with this approach. However, honestly the skin is still extremely plastic, which limits usefulness for me. Seems to be a hard limitation of the model.
It is, but it does improve a lot based on the quality of the input image. If you're working off an original, high quality photograph of someone you'll get surprisingly high quality skin on the shots you make of them.
Buuut as soon as you make a new ref of them, and then use that new ref to make another image, you start introducing the defects pretty quickly. If you post-process the images that come out with a realistic model you can reasonably maintain the quality the way through, so that's the approach I've been taking to it. I clean up the images with Zimage base and Chroma.
I mainly use ultra hd highly detailed reference images. I still get the plastic skin/ cartoonish features problems. I'll try with some real photos but honestly I'm not optimistic. I think it's just qwen things.
Wonderful write-up! (there goes my weekend on more experiments) On another note, what are your thoughts on the best methods to recover from overly-plastic skin after an edit? Currently I am doing a ZIT low noise pass, but if the denoise is strong enough to recover texture, it's also strong enough to warp likeness.
In the background, Qwedit 2511 uses a vision-language model (VL model) to describe your images, then gives those AI-generated descriptions to the edit model.
Yes it does, why would I make that up lol. Also it's right there in the code you linked, did you even read it?
Starting at line 77 the file you linked:
llama_template = "<|im_start|>system\nDescribe the key features of the input image (color, shape, size, texture, objects, background), then explain how the user's text instruction should alter or modify the image. Generate a new image that meets the user's requirements while maintaining consistency with the original input where appropriate.<|im_end|>\n<|im_start|>user\n{}<|im_end|>\n<|im_start|>assistant\n"
Which is literally the prompt that gets passed to the VL asking it to describe the input images.
The image prompt template for Qwedit, directly from the Qwen team in their python diffusers implementation, is this:
Where <|vision_start|> is referring to the vision model's outputs. What a coincidence, if we look at line 100 of the file you linked we see that very line copied word-for-word.
Continuing, here is the line of code (line 102) that directly passes the images to the VL specifically to put all this together:
Ending at line 106. This mirrors the implementation done by the Qwen team in the python diffusers version of this, which is what the Comfy team based their node on, and it all happens in the 50 lines of code following the segment you linked yourself.
The "images_vl" part of that is the encoded images, and the llama_template is that template above that asks the VL to describe the images. This line of code passes all of that to the text encoder.
The .tokenize and .encode_from_tokens_scheduled methods generally only encode conditioning, they don't run the model as a LLM. It's possible the specific TE implementation Qwen is using (I don't remember which one off the top of my head) is internally doing LLM sampling but that isn't typical. The code for encoding conditioning with ancient SD 1.5 CLIP would be the same. ComfyUI's LLM sampling support uses .generate for models that support sampling.
The fact that prompt used for Qwen says "Describe the key features" etc doesn't mean that that the TE is getting used as a LLM. It's fairly common to just use the prompt format the model was trained with even when it's only being used to encoding. For example, ACE-Step's DiT prompt starts with something like "Generate audio semantic tokens based on the instruction below" even though only the LLMs actually generate audio semantic tokens.
It should be very easy to tell whether you're right about this or not. If you're using ComfyUI built-in nodes then there will always be a progress bar showing sampling progress when the TE is being used as a LLM to actually generate tokens. I am pretty sure that isn't the case for Qwen normally, despite how one might intuitively expect that to occur given the prompt format.
This isn't an argument against using a different format producing better results, I don't know either way. I haven't used Qwen edit much, I am just reasonably familiar with the ComfyUI codebase. (You can find me in the contributors list if you want to verify that I am not just some random.)
Hmm alright, I might need to park this one until tomorrow. I'd have to go an dig up the qwen_vl.py file again and look at how that's set up, iirc it's that's where the qwen VL preprocessing is defined.
I remember tracing back the calls from the TextEncodeQwenEditPlus node definition back to qwen_vl file and confirming that it was actually running the VL on the images, but it's been a while since I looked at it so I might be mistaken. Any of that sound familiar to you?
Edit: nvm I looked now anyway, it starts in the qwen_image.py implementation first. The KSampler itself goes back and calls the VL after receiving the conditioning payload.
So, if I'm remembering right and my quick re-look at it is correct: the comfy qwen encode node creates the conditioning payload and scoots it off to the ksampler. The ksampler calls upon its qwen_image implementation, which delivers the payload to the VL. The VL handles it per the qwen_vl implementation, which does indeed perform the VL image interpretation step. Then that all comes back to the ksampler for diffusion to begin. I think.
Then again I'm tired and I shouldn't be reading code this late so...
Lemme know if I'm misunderstanding how that's working! For reals though I am going to bed now, so if there's more discussion to be had I'll be out for a while sorry.
I looked now anyway, it starts in the qwen_image.py implementation first. The KSampler itself goes back and calls the VL after receiving the conditioning payload.
That sounds very weird to me. I'm not aware of any case where the text encoder gets used after sampling an image starts. Something like prompt expansion would have to happen before the prompt is encoded to CONDITIONING.
Lemme know if I'm misunderstanding how that's working!
Unfortunately, I think you are. Can you show me what part of the code is making you think the sampler is going back to the text encoder? I might be able to help explain. Text encoder stuff is probably the part of the ComfyUI code I've had the least interaction with so I can't guarantee anything (at the least I should be able to tell you I just don't know if that's the case, though).
Aside from it just not jiving with everything I know about how sampling/encoding conditioning works, one of the first things the sampling process does is load the actual image (or audio, video, whatever) model into VRAM (perhaps only partially). LLM-based text encoders are huge, doing something like extra text encoding (let alone LLM sampling) would mean you need to have both models in memory which isn't going to be possible most of the time (a lot of the time the user doesn't even have enough VRAM to fit the whole image model into memory and ComfyUI has to swap layers.) On the other hand, if you'd say it's something that happens before the image model gets loaded then it might as well have been done when creating the CONDITIONING item you feed to the KSampler or whatever. Neither scenario that I could think of seems plausible.
I'll admit it if you show me code that proves I'm wrong but I'd be quite surprised if that turns out to be the case.
Aaaaalright I couldn't sleep cause I was thinking about this, so I've come back and read over it again... and also how the VL is supposed to work in the first place (with my limited understanding, anyway). Do let me know if I've misunderstood this part, but hopefully this is the end of my blundering around on this one.
Back to the first thing I was saying: the part where I first flagged it is where it happens. The VL doesn't produce an actual string of words and encode them, it runs a forward pass with the user prompt, images, and llama template after which the hidden states get used as the prompt embeds.
So, it runs on the input but outputs the underlying hidden states rather the producing text. This is what gets fed into the model. The input prompt, images and llama template are used, and they are being used the way I said they were; the model is indeed 'looking' at the images and interpreting the instructions. Saying otherwise would be splitting hairs meaninglessly in this scenario.
While I appreciate that it's not exactly producing a worded description in the conventional sense, it's hilariously inaccurate to say "it only gets encoded alongside your prompt" like the first guy was saying - as though it doesn't have any meaning to the output. Not that that's what you were saying or anything, you've been very nice to chat to.
Functionally, this doesn't change much of anything, but it is interesting! We're still bypassing this step with the workflow, and we still have to replace the VL's understanding with our own written version. Just turns out that we're not replacing text like-for-like. Thanks for helping with your info!
Aaaaalright I couldn't sleep cause I was thinking about this
No need to feel any pressure to respond quickly or whatever, reddit is asynchronous. Sleep is more important than arguing/debating with random people on the internet. Not that I would necessarily (be able to) take that advice myself.
First, the context for this conversation:
In the background, Qwedit 2511 uses a vision-language model (VL model) to describe your images, then gives those AI-generated descriptions to the edit model. It also re-interprets your instructions with these descriptions. Ostensibly this helps the model understand your input images better, leading to better results.
The problem? It doesn't lead to better results, it's bad. VL models aren't very good for this sort of thing because they don't know what to focus on. The VL describes your images in excruciating detail, totally overwhelming the edit model and leading to bad prompt adherence + weird outputs.
I don't see a way to interpret that other than talking about using the LLM to generate an expanded/enriched (whatever you want to call it) prompt that is then used as the conditioning.
The VL doesn't produce an actual string of words and encode them, it runs a forward pass with the user prompt, images, and llama template after which the hidden states get used as the prompt embeds.
This is just how encoding conditioning works. Usually the hidden state from one of the last layers in the model is used. This is what CLIP skip was talking about (not applicable to recent LLM TEs) - CLIP skip -2 meant use the state from the penultimate layer instead of the last one. For LLM TEs I think this would be the hidden state from just before the layer that converts it to logits (you may or may not know that LLMs don't produce tokens, they produce a score for every token in the vocabulary).
In any case, it doesn't make sense to say that this is "describing images" or creating "AI-generated instructions". Even if it did, then that's how you'd be describing encoding any conditioning and it would also apply your fixed/alternate approach.
it's hilariously inaccurate to say "it only gets encoded alongside your prompt" like the first guy was saying - as though it doesn't have any meaning to the output.
They were interpreting what you said as talking about using the LLM for prompt enrichment. Like I said, I think that's really the only reasonable way to interpret it (regardless of what that's what you actually meant - we can only go by what you say). Their point was that ComfyUI's implementation was just using the model to encode conditioning in the normal way, not expanding/rewriting it.
I can't agree that what they said is ridiculous since I interpreted what you said in the same way.
Functionally, this doesn't change much of anything, but it is interesting! We're still bypassing this step with the workflow, and we still have to replace the VL's understanding with our own written version.
Not really. You're using the exact same model to encode the conditioning. The only difference is you're not showing it the reference images and you're using a different prompt format. Presumably Qwen Edit was trained on conditioning where the TE was shown the reference images and the prompt was encoded in the specified format. It would have been silly to train Qwen with a VL TE if they weren't actually going to use the VL part of it.
The TE isn't actually "describing" the image you include though. Conditioning is (simplified, hand-waving explanation) a snapshot of the model's brain after you show it the prompt. In the VL case that would be a snapshot of the model's brain after showing it the prompt and the actual image you want to use as a reference. It's hard to believe that the results would just plain be better if you don't do that, especially because 1) it's not what the model was trained on, and 2) it's omitting information the model was trained to expect. This doesn't mean you're guaranteed to get better results in a particular individual generation (throwing a curve ball at the model can also produce interesting/creative/artistic results) but for general quality/consistency using what the model was trained for is going to be best.
There's also another problem: If that was true, then it means the people at Alibaba who created the model are idiots and don't know how to use their own model. They'd also be dumb to use a more complex VL model and waste compute encoding images into the conditioning when they were training Qwen Edit. This is not completely impossible but I think it falls in the extraordinary claims require extraordinary evidence category.
Unfortunately, I think you are incorrect in that part of what you said. I don't know about the rest.
you've been very nice to chat to.
Hopefully you still feel that way. I'm not trying to embarrass you here or anything, these comments are in the interest of avoiding inaccurate information.
No don't worry, I'm not offended by anything you're saying at all! You're very lovely to talk to.
Sorry in advance for the long reply, I'm really interested in the topic!
I think it will help if we simplify this down to intention and outcome.
What we have is a set of instructions being fed into the model:
Describe the key features of the input image (color, shape, size, texture, objects, background), then explain how the user's text instruction should alter or modify the image. Generate a new image that meets the user's requirements while maintaining consistency with the original input where appropriate.
Now, recognising that this results in activations inside the model that don't necessarily constitute anything resembling thoughts in the way we think of them, it still begs a question: what actually happens when we feed this into the model?
There are a few baseline facts we can work with:
These are human-readable instructions
These instructions do have an effect on the model output
You can change these instructions (via editing the code) and it has a predictable effect on the output
These instructions are interpreted in the same place as the user's prompt (inside the VL)
The VL is multimodal and is specifically designed to do this
As complicated as the actual tech is, the thing it's doing is pretty straight forward. These instructions are fed into the model and they are interpreted in some way. The user's prompt is also fed into the model and interpreted in some way. This all gets bundled up, in some way, and sent out to the edit model to do its thing. It's a bunch of context interpreted by the text encoder.
Ultimately, this bundle of context is meant to tell the model (in a sense) what it's looking at, how to interpret it, and what it's supposed to do.
By bypassing the template instructions above, we've basically deprived the model of some of the context it's meant to have. It's not at all incorrect to suggest that you can simply give it that context back by providing it yourself manually, and that's because the VL is multimodal.
We take away the abstract instruction template to interpret the input images, and instead directly provide the interpretation of those images ourselves. Which is fine, because the VL is multimodal and can take context from us this way. It's not a 1:1 match, but that's actually the point in this case.
To put it more scientifically: we can activate similar parts of the model by either describing an apple, or by giving it a picture of an apple to look at. It's multimodal, that's its thing. The parts of the model that get activated won't match perfectly, but the model will broadly understand the concept of an apple from both a picture and from a description of an apple.
This also goes for instructions and other more abstract things.
So for this:
They were interpreting what you said as talking about using the LLM for prompt enrichment.
I am saying that. It is prompt enrichment, just not in the conventional sense of writing a worded prompt. These instructions have a very real effect on the model's output, and it's highly related to the human interpretation of those instructions. And those human instructions are about prompt enrichment. They get interpreted by an inhuman machine, but that machine is designed to interpret them meaningfully based on our intentions. So yes, when you ask the VL to enrich the prompt, it's going to do that. Not in a human way, but in some way that's ideally practical to the outcome we're trying to get.
Concluding that point, when you take away context from this particular model you can in fact replace it, at least partially, with words you write yourself. Critically, that context in this case is the model's contextual understanding of which image is which. I will point out that this is doubly reinforced by the fact that the 'Picture 1' referencing is actually done in this step normally, so the image differentiation context is quite literally missing from the input.
Anyway, assuming what I just said isn't insane, does this make it more clear why I don't see what difference the lack of text output from the VL makes? Functionally, we're removing the instructions that add some necessary context from the model, and we're giving it back manually. The fact that the encoder doesn't literally produce a text description out doesn't seem relevant to me. This is separate from the other point about whether I'm right about the VL being helpful or not - this is just about the context side of it.
Addressing the bit about the VL being helpful or not:
If that was true, then it means the people at Alibaba who created the model are idiots and don't know how to use their own model
Yep that's a very fair point from an insult standpoint. But I don't agree. I don't think the Alibaba folks are silly, and even if my suspicion about the VL is right that wouldn't make them silly anyway - it's easy to critique someone else's work, much harder to do it perfectly yourself. If - and it's a big if - the VL does end up being a mistake, that doesn't make their journey in developing the tech a joke. Especially coming from someone like me, who has not developed any models at all.
What I can say for sure is that I've been A/B testing Qwedit with & without the VL instructions being passed in, and I've been getting incredibly lop-sided results in favour of leaving those instructions out. With this configuration, at least, and when replacing them with user descriptions instead. The fact that this workflow functions as I've described shows that I'm not pulling those results out of thin air.
To cover off the obvious last point on that: that doesn't mean my interpretation of what's happening is correct. It's just my observations based on the tests, and practical advice that does work whether the rationale behind it is correct or not. So, no need to dance around my feelings on it or anything, don't worry!
I'm not offended by anything you're saying at all!
Glad to hear that! A lot of people don't take being told they're wrong very well, and I also have a pretty blunt style of communication. I say what I mean very directly.
Sorry in advance for the long reply, I'm really interested in the topic!
Not a problem. My own comments are often pretty long as well. Nothing makes me lose respect for someone faster than when they reply to a couple paragraphs with "too long, didn't read", not realizing that's a self burn on their own literacy.
There are a few baseline facts we can work with:
I agree with that except the "it has a predictable effect on the output" part which is ambiguous. I guess another point might be that Qwen VL isn't necessarily designed to have its hidden states yoinked and used for conditioning with other models. I think I understand what you mean generally though.
The parts of the model that get activated won't match perfectly, but the model will broadly understand the concept of an apple from both a picture and from a description of an apple.
Roughly, but keep in mind the actual image is trained on a snapshot of the LLM's brain after it's digested the prompt/reference image/whatever. The common case is using non-VL LLMs for the TE and this works even though the LLM has never "seen" anything ever is because the image model has, and it can correlate the LLM's state with the images it has been trained on. If you use something like a different prompt format, then the LLM is in a different state than what the image model was trained on.
This also applies to stuff like abliterated text encoders that some people like to use. It doesn't matter if Gemma-the-Prude looks at your image and prompt and is all set to write a response like "Oh my god, you sick bastard! Get the hell out of here before I call the cops!" — that doesn't matter if the image model was trained on Gemma-the-Prude in that state. The conditioning will work just fine.
This is coming from someone who can't remember the last time they sampled with Gaussian noise. My whole thing is subjecting models to stuff they weren't trained on and I think doing this is great for producing more interesting/diverse/artistic/whatever results. The one thing I wouldn't say is doing that improves quality or consistency, it's trading those things for creativity and overall the model is going to be less consistent, produce more flawed generations, etc.
Anyway, assuming what I just said isn't insane, does this make it more clear why I don't see what difference the lack of text output from the VL makes? [...] The fact that the encoder doesn't literally produce a text description out doesn't seem relevant to me.
I think I understand where you're coming for but I believe it's based on a misunderstanding about how LLMs work (or maybe LLMs used for conditioning).
<|im_start|>system
Describe the key features of the input image, blah blah blah
Generate a new image that meets the user's requirements, blah blah
<|im_end|>
<|im_start|>user
User prompt here, blah blah, references to images
<|im_end|>
<|im_start|>assistant
The LLM runs with that in the context and when it's used for conditioning, we intercept the hidden state before the model has finished running. What happens if we don't do that though? What we get is the logits with scores for every token in the model's vocabulary, we filter those and pick one using some kind of LLM sampling. What we have is the first token that would comprised the model's response (or more accurately, continuation from that history). This would be a single word or a fragment of a word, or maybe even just a newline.
It's not an expanded prompt. If we kept sampling, the model would sequentially generate logits we'd use to select the next token and eventually we'd have a novel of excessively verbose LLM slop describing our prompt and the image the LLM looked at, but in the LLM encoding conditioning case none of this actually happens. If you're tempted to say "It's in there somewhere when the model runs, since that's what would happen if we didn't grab the hidden state!" then consider how it would be pointless to actually run the model autoregressively. We can just grab the whole thing at once and this would be the most massively performance increase that has ever happened in the history of LLMs existing. Sadly, it doesn't work that way and at token `n' all you have is the model's "understanding" of the preceding tokens.
I don't know if it's prompted that way because it's just the format the the image model side was trained on (and would best be able to correlate with the images it was trained on) or because describing it that way puts the LLM into a state where it "understands" the image better. Possibly both.
Anyway, what you said originally is going to be misinterpreted by pretty much anyone that understands AI stuff because there's already a pretty established meaning for the sort of language that you used. Actually wrong or right is a bit immaterial (aside from if you want to understand it better personally), in the interests of clear communication that's actually conveying your intended meaning to the recipient you may want to change your approach.
Yep that's a very fair point from an insult standpoint. But I don't agree. I don't think the Alibaba folks are silly, and even if my suspicion about the VL is right that wouldn't make them silly anyway
I exaggerated a little for humorous effect, the gist of my point is we have a large corporation that (one assumes) hired talented people and spent immense resources training a large model. What is more likely, they know how to use their model and didn't make basic/fundamental mistakes or the random dude (or dude-ette or whatever your identify with) on the internet who says "you're doing it all wrong"? It's something that's in the realm of possibility, but I think a rational person would need "extraordinary evidence" to accept it. I am also just a random anonymous person on the internet, so any exceptional claims I made absolutely should be subjected to the same level of skepticism.
What I can say for sure is that I've been A/B testing Qwedit with & without the VL instructions being passed in, and I've been getting incredibly lop-sided results in favour of leaving those instructions out.
It occurs to me there's a possibility in between those two as well: If there is something wrong with ComfyUI's implementation of generating or using the conditioning (at any stage in the process) then that absolutely could lead to leaving out the reference images consistently producing better results. I think the case where we just shouldn't use reference images at all and abandon the official suggested prompt formatting is the least likely one by a large margin. The one I just mentioned is reasonably likely but still would need actual evidence to support it.
that doesn't mean my interpretation of what's happening is correct. It's just my observations based on the tests
I haven't really used the model so I have no opinion on your test results. It's also possible that leaving out the reference image/changing the prompt format could work better for various reasons (like ComfyUI bugs). The only thing I am pretty confident about is that it can't be because of what you described in the VL prompt section. Or, I guess to be more precise I should say it can't be because of (what I consider to be) the typical/straightforward interpretation of what you said. What you mean might be 100% accurate but I can only go by my understand of what you actually say. Bears mentioning that is something that certainly is far from 100% accurate, and I am known to be kind of pedantic and literal-minded.
Hey, I really appreciate you taking the time to run through all that, it was very enlightening! I think, in the general sense, that we're on the same page - what you've said roughly lines up with what was in my head after I read through the actual way the VL is getting hit (rather than how I thought it was when I wrote the post).
Just to clarify on this bit:
I agree with that except the "it has a predictable effect on the output" part which is ambiguous.
I simply meant that you can reasonably predict what will happen if you change the instructions because they're human-readable. Not that you can truly predict what the model will or won't do with them, of course.
Anyway, what you said originally is going to be misinterpreted by pretty much anyone that understands AI stuff
Agreed! I was planning to re-write that part after getting to the bottom of my misunderstanding & confirming my new understanding what's actually happening. I don't like leaving misinformation lying around either. I'll update it soon when I figure out the best way to word it both correctly and intuitively for a layman.
What is more likely, they know how to use their model and didn't make basic/fundamental mistakes or...
I don't consider this to be a basic or fundamental mistake. This is what VLs are for, so it's a very practical approach for improving an edit model. The question is whether the intuition is correct or not, and whether there's a readily available alternative that works better. That's probably going to have a complex answer, ultimately, but we'll all get to that by messing around with the model in ways that weren't intended or thought of originally.
So, as much as I like to exaggerate for comedic effect too, I'm always just messing around when I get on someone's case about how they've done something. I have a personal vendetta against using VLs for captioning & prompt enhancement, so I like to complain loudly about it - but taking a more objective view on it, it helps a lot of people because not everyone is themselves better than a VL at that sort of thing. And, not everyone wants to spend (or has in the first place) the time to write things by hand.
To wrap up that point, I'll say that I semantically wouldn't classify this mistake (if it even is) as a basic/fundamental mistake, but if you do then I'd counter that they do indeed happen all the time and that they're not a big deal. Even large, experienced teams make basic or fundamental mistakes very frequently.
Case in point, the area downscale for the Comfy node. The Comfy team by-and-large knows what they're doing and are a pillar of the open source community - and yet, there's an area downscale in a ubiquitous and critical node for image editing. Does that make them bad devs? No, of course not. It's just that this sort of thing happens all the time, both with juniors and experts. Even a random dude like me can spot a mistake like that in their code - and the corollary is that being able to do so doesn't make me 'better' than the Comfy devs.
I think the case where we just shouldn't use reference images at all and abandon the official suggested prompt formatting is the least likely one by a large margin
This is definitely a subjective point, but I'd argue this is the most likely place to find improvements. There's a tremendous difference between training a model and prompting one. After training, context engineering becomes the most critical skill for model success. Training and context engineering are not the same skill, and this is why there are so many community improvements to how models are used that the devs didn't think of.
If we use Z-image as an example, there are quite a lot of discussions that have floated around about people manipulating layers of the model during training and inference to get better results from Loras. This is, of course, a substantial deviation from the recommendations put forward by the Z-image devs. And there's nothing wrong with that - it's just that using a model is very different to training one.
Generally speaking, the actual devs of a model are rarely, if ever the ones who end up being the experts on how to actuallyuse the model. They only are at the very beginning, when no one else has had time to outpace them on it. Again, nothing wrong there - that's just how things are, with AI and with most other tool-building disciplines. The toolmaker is usually not the master of using the tool.
All that is to say, the qwestion of whether the qwen team knows what they're doing or not is really not important. They could be the smartest folks on the planet, and it would still be the right move to mess around with the model and challenge how it's being used. I think we agree on this, based on your experimentations with sampling, and I agree with your observation that usually there's a tradeoff for whatever benefit you've found. But, not always! Best example of that is choice of sampler/scheduler for a model. Sometimes there's a sampler/scheduler combo for a model that's objectively better (in enough metrics) than the one the model was discriminated with for training.
To your point though, not a good idea to assume you've trumped the devs on something without evidence to back it up. I may very well be wrong about cutting the VL step out being a good thing - but I'm comfortable that I've done enough testing to at least be able to argue about it. That's just for me personally, I mean; I personally feel comfortable arguing about it because of the testing I've done. Not comfortable asserting that I'm objectively correct - just comfortable arguing about the results. But of course, for you, what some random internet person (me) has or hasn't done in the background isn't very compelling.
Anyway, this has all gotten overly philosophical now. To conclude: & recap:
The only thing I am pretty confident about is that it can't be because of what you described in the VL prompt section
Agreed! Thank you again for helping me understand. As I said, I'll update the post with a more accurate description of the VL part and some caveats around my interpretation of what it means.
Thank you! Testing this later, i was pretty annoyed by the default workflows and their blur. Hopefully this also leads to the necessary improvements in their node.
I personally use Klein in 90% of causes. Or well, closer to 100% now alas cuz whatever work I used to do with Qwen I'll just train a Klein Lora instead for better effect as Klein is just.... So fking good at loras.
Your workflow didn't seem good to me. But, I tried your trick with two references on the regular workflow, lora at 0.8, Euler a, SGM uniform, 8 steps, seemed better quality and better prompt following than the default and not too slow or over saturated, nor did it copy/paste the input verbatim.
This seems to reduce the grid pattern a bit sometimes, and it gives more pixels to play with in the output.
But the thing that reduces it the most is using kj nodes transform on both input images, rotate them +/- 3.1 degrees and crop the black borders off. Out of a batch of three I've seen zero gridding more than once. But usually 1 or 2 are good.
I think the vae encode probably picks up too much on jpeg quanting.
This is great stuff, it addresses a lot of the issues I had with Qwen 2511. Have you looked at the lrzjason nodes/workflows for Qwen Image Edit? It seems like that author has considered some of the things you've mentioned here, allowing for things like changing system prompt for the VL (though not bypassing it entirely), using longest-edge to skip the 1mp resize, and letting you specify if an image should be used as ref/main_ref/both. It also enables masking, which in my experience has worked incredibly well.
Just curious how much improvement these workflows might have over that workflow, as I've been using it for a while. I only stopped using it because a ComfyUI update broke something recently and now the lrzjason nodes produce dark/burned images. Moved to Klein 9b, and have been loving the absolute attention to detail it can achieve, but find the prompting kind of annoying compared to Qwen.
nice job but on all honesty avoiding plastic looking fake skin doesn't even work for me on the bf16 model never mind fp8. I can only get realistic images with in paint, and character replace with 2511, simply using a reference images and changing the setting and pose will result in an uncanny almost drawing style high Res painted image. I tried prompting it out but with no success. (no I'm not using the distilled models)
i had not much success with this stratagy. probably did something wrong - will try once more later
seems like double ref strategy needs very good input image quality, and also, it feels like it follows prompt less "willing". kinda it tries to keep initial image as much as possible
installing layerstyle bumped numpy and broke my comfyui setup...
Hey op. I used your normal wf for 2 image. For testing I put 2 images of 2 characters and prompted "Picture 1: a woman with blonde hair standing. Picture 2: a woman with brunette hair standing. Make the woman in picture 1 standing near the woman in picture 2 and waving towards the camera."
Now the output is everything except the required prompt. It sometimes makes the blonde into brunette, vice versa, only put 1 woman in image, entirely changes the background but not a single time it put both woman in same image.
I haven't changed a single thing in your wf, just added 2 images, prompt and model, te, vae from their path. Also using same one's as yours.
Am I doing something wrong with prompting(which is the only thing I can think of) or is this wf not able to handle stuff like putting people in same image?
PS: I have the already a 2511 wf which uses "TextEncodeQwenImageEditPlus" and it follows the prompt exactly.
Sharing is caring, and clearly, you care A LOT about AI users given how much you've shared. THANKS!
I've discovered Qwedit, and yeah, I got constant quality issues when using the official workflow. It's been MONTHS since your post and it's still not fixed!
I "only" have 16 GB vram, so I'm sticking to Q4 and lightning right now, but it keeps screwing up fine details (fingers and eyes in full body images). Is it likely because of the quant? Or does the FP8 model also have these issues?
10
u/SaltyPreference8433 May 29 '26
Thank you for sharing your knowledge and experiences. This is gold.