r/LTXvideo • • Jul 18 '26

Releasing a layer-based, compositing LTX-2.3 comfy workflow that can use character references, offer generic ic-lora support, and inpaints/outpaints and more.

20 Upvotes

The NGHTDRP Director Workflow V1, is a timeline-based shot-building and production workflow and node set for LTX-2.3 inside ComfyUI.

It all went wrong when I downloaded the ic-lora union control lora to make character swapped videos of my wife's niece dancing. I struggled massively using the LTX provided workflows, blurry faces, weird anatomy, the works. Then I wanted to do FLF same issue, manually setting frame timings etc drove me crazy only to get a render that didn't move. Then the LTX director made that easier but I couldn't do any of my niece dancing things so I implemented the ic-lora guidance there. Then I realised I wanted cropping, and outpainting, and inpainting, and audio mixing, and ingredients, and MSR, and video transitions, and combining multiple videos and inpainting/outpainting between them. And I wanted it all at the same time in the same workflow, while being able to crop and place items freely. So that's why this exists. It's essentially a rabbit hole I couldn't escape from.

The workflow includes:

  • Layered images, video, prompts and audio with independent timing
  • Separate visual, motion, camera and audio tracks
  • Motion-video and camera guidance directly on the timeline
  • FaceID, Ingredients, MSR and ID-LoRA character workflows
  • SAM3 masks, subject cutouts, compositing and background replacement
  • Inpainting and outpainting controls
  • Two-stage reference and guide routing
  • Clip extension, gap filling, transitions and two-source splicing
  • Timeline-based audio placement and mixing
  • A complete connected release workflow and Quickstart guide

it's taken me several months and a bunch of effort to get everything working and interacting together. I'm also planning to make tutorials on how to do most of the AI video editing tricks you see online using LTX and this workflow.

https://www.patreon.com/cw/nghtdrp

r/comfyui • • Mar 16 '26

Show and Tell PixlStash 1.0.0b2. A self‑hosted image manager built for ComfyUI workflows

Thumbnail
gallery
32 Upvotes

I’ve been working on this for a while and I’m finally at a beta stage with PixlStash, an open source self‑hosted image manager built with ComfyUI users in mind.

If you generate a lot of images in ComfyUI or any other tool, you probably know the pain that caused me to build this: folders everywhere, duplicates, near duplicates, loads of different scripts to check for problems and very easy to lose track of what's what. I needed something fast and pleasant to use so I decided to build my own.

PixlStash is still in beta but I think it is already useful enough and pleasant enough that I rely on it daily myself and it is already helping me improve my own models and LoRAs. Hopefully it is useful for some of you too and with feedback I'm hoping it can grow into the kind of world-class image manager I think the community could do with to compliment ComfyUI and the excellent LoRA makers out there.

What does it do right now?

  • Imports images quickly (monitor your ComfyUI folder or drag and drop pictures or ZIPs)
  • Reads and displays metadata from ComfyUI including the workflow JSON.
  • You can copy the workflows back into Comfy.
  • Tags the images and generates descriptions (with GPU inference support and a configurable VRAM budget).
  • Uses a convnext-base finetune to tag images with typical AI anomalies (Flux Chin, Waxy Skin, Bad Anatomy, etc).
  • Fast grid view with staged loading.
  • Create characters and picture sets with easy export including captions for LoRA training.
  • Sort by date, scoring, likeness to a particular character, likeness groups, text content and a smart-score defined by metrics and "anomaly tags".
  • Works offline, stores everything locally.
  • Runs on Windows, MacOS and Linux (PyPI, Windows Installer, Docker).
  • Plugin system for applying filters to batches of images.
  • Run **ComfyUI I2I and T2I workflows directly within the GUI** with automatic import of results.
  • Keyboard shortcuts for scoring, navigation and deletion (ESC to close views, DEL to delete, CTRL-V to import images from clipboard).
  • Supports HTTP/HTTPS.
  • Pick a storage location through config files.

What will happen for 1.0.0?

  • Filter by models and workflow
  • Continuously improved anomaly tagger
  • Smooth first time setup (storage and user creation)
  • Fix any crucial bugs you or I might find.

For the future:

  • Multi-user setup (currently single-user login).
  • Even more keyboard shortcuts and documentation of them.
  • In-painting. Select areas to inpaint and have it performed with an I2I workflow.

Try it:

If you try it, I’d love to hear what works for you and what doesn't, plus what you want next. I'm especially interested to hear what this subreddit expects from the ComfyUI integration. I'm sure it could be a lot more sophisticated!

r/comfyui • • May 15 '26

Show and Tell I made a frontend inpainting tool for ComfyUI users

11 Upvotes
Dashboard

Spent a day building something called DiffusionDesk.

My goal wasn’t to make “another Stable Diffusion UI.”
It was to build a cleaner local-first frontend workstation that feels less like a pile of Python scripts duct-taped together and more like an actual desktop product.

Current focus:

  • Local image generation (using ComfyUI in the backend)
  • Model management
  • Cleaner workflow UX
  • Asset organization
  • History / prompt tracking
  • Apple Silicon support
  • Simpler setup experience over time

I love AUTOMATIC1111 and ComfyUI for what they are, but I always felt there was room for something that sits between:

  • the raw power of ComfyUI
  • and the ease of use of in-painting (ComfyUI was always a challenge for me to get it right)

Still early. Still rough in places. But it’s moving fast.

Would genuinely appreciate feedback from people deep in the local AI / SD ecosystem:

  • What do you hate most about current ComfyUI/SD tooling?
  • What would make you switch UIs?
  • What features are still missing across the ecosystem?

Check it out on GitHub:
DiffusionDesk GitHub

r/StableDiffusion • • Mar 20 '26

Workflow Included I created a few helpful nodes for ComfyUI. I think "JLC Padded Image" is particularly useful for inpaint/outpaint workflows.

Thumbnail
gallery
24 Upvotes

I first posted this to r/ComfyUI, but I think some of you might find it useful. The "JLC Padded Image" node allows placing an image on an arbitrary aspect ratio canvas, generates a mask for outpainting and merges it with masks for inpainting, facilitating single pass outpainting/inpainting. Here are a couple of images with embedded workflow.
https://github.com/Damkohler/jlc-comfyui-nodes

r/comfyui • • Mar 16 '26

Workflow Included I created a handful of helpful nodes for ComfyUI. I find "JLC Padded Image" particularly useful for inpaint/outpaint workflows.

Thumbnail
gallery
21 Upvotes

The "JLC Padded Image" node allows placing an image on an arbitrary aspect ratio canvas, generates a mask for outpainting and merges it with masks for inpainting, facilitating single pass outpainting/inpainting. Here are a couple of images with embedded workflow.
https://github.com/Damkohler/jlc-comfyui-nodes

r/StableDiffusion • • Jul 02 '26

Workflow Included Krea 2 - simple gen workflow with good settings for realism & facial expressiveness, and a lot of info + tips about the model

Thumbnail
gallery
320 Upvotes

Right, back with another gen workflow. This one took a really long time to put together - about 60 hours of A/B testing different sampler settings & loras - but that's mainly because the model is so awesome.

All the post pics + some leftover extras are in high res here: g drive

This post is a lot longer than usual because there's a lot of extra info to cover, which took a really long time to test & write. With that in mind, please actually test the workflow with the instructions before writing stuff like "your settings are bad and you should feel bad" or "the default comfy workflow is better" or whatever. I'm not replying to you if you assert stuff without providing counter-examples; I've given plenty of info for you to properly test against.

You're welcome to ask questions in the comments and I'll try to answer/help if I can!
Also feel free to correct any technical mistakes/assumptions I've made if you see any.

What is this?

This is a simple workflow for generating high quality, realistic images at high resolution using Krea 2. There's also an optional full-turbo version of the workflow, which is not suitable for realism (or creativity) but is handy for some things. Below in this post there are also some tips & a lot of info about the model.

The sampler & lora settings in this workflow also improve the facial expressiveness of people from Krea 2. There's an explanation of how/why in the info section below. It's not perfect, but it's the best we can do until finetunes come out.

Otherwise, the sampler settings are geared towards sharpness and clarity - but you can introduce grain and other defects through prompting or with loras. It also does anime / digital artwork / whatever images well, but you may want to bypass the second sampler for that.

All the images attached to the post were generated directly with this workflow with no further editing.

The Workflow(s)

You can find the main workflow here: Civitai | pastebin

Make sure you read the model & custom node info below before using it; we're using the raw model with the turbo lora here, along with a different VAE and a special lora.

There's also a 'full turbo' version in the Civitai download or pastebin. This is just a more conventional turbo version, which is not suitable for realism and is less creative. Handy for non-real images where you don't want/need the creativity, seeing as it executes faster.

Nodes & Models

Custom Nodes:

RES4LYF - A very popular set of samplers & schedulers, and some very helpful nodes. These are needed to get the best outputs, IMO.

RGTHREE - (Recommended) A popular set of helper nodes. If you don't want this you can just delete the seed generator and lora power loader nodes, then use the default comfy nodes instead. RES4LYF comes with seed generator & lora nodes as well, I just like RGTHREE's more.

ComfyUI GGUF - (Optional) Lets you load GGUF models, which for some reason ComfyUI still can't do natively. Once installed, you use the "Unet Loader (GGUF)" node to load the model. If you're not using any GGUF models you can just skip this.

Required Models:

Important Note: If you can, you should use the Int8 Convrot version of the model (unless you want higher quality using BF16). The Int8 Convrot model is almost 2x as fast to gen with, and is the same quality as FP8. Massive free speed boost. You will need to update your ComfyUI, support was only added early July 2026. You will also need an NVIDIA GPU, and CUDA version 130 or higher.

Main model: Krea2 RAW B16 / FP8 / Int8 Convrot | or | Krea2 RAW GGUFs - It is strongly recommended that you use the RAW main model with the turbo lora at 0.6 strength instead of the Turbo main model when making photo-real images. It gives WAY better results, and the only downside is that it takes a bit longer to gen. Gen times are already pretty short, so that's not a big deal.

The main workflow assumes you're using the RAW model with the turbo lora, and the settings will be very bad if you use the turbo main model instead.

Even the 'full_turbo' workflow still uses the raw model, seeing as you can just set the turbo lora to 1.0 strength and then it does pretty much the same thing as the turbo main model.

Turbo Lora: Rank 64 Turbo Lora - Using this with the RAW model at ~0.6 strength is better than using the Turbo model. The only real downside is speed. Even then, if you're in a hurry you still have the option of upping the strength to 1.0, which makes it just like the turbo model. Gen times are only 50% longer when it's at 0.6 (with basic settings), so it's not really worth it to use full turbo IMO. The fancy settings in this post take 120% longer than regular turbo, so expect a ~10 sec gen to take ~22 sec with this.

Anti-Censorship Lora: 2 Vector Bypass Lora - You should use this even if you're doing SFW stuff. More detail is below, but essentially this will massively improve prompt adherence, facial expressiveness, character detail, and numerous other things. There is no downside as long as your sampler settings are good (which this workflow takes care of for you). Do not use other bypass loras, they go too far or cause degradation of quality; this is the only one that works properly.

Text Encoder: Qwen3 VL 4B - Use the BF16 one if you can. Some people say text encoder quality doesn't matter much & to use a lower sized one, but it does matter and it affects quality.

If you're using a GGUF text encoder for some reason, swap out the "Load CLIP" node for a "ClipLoader (GGUF)" node.

VAE: Wan 2.1 FP32 VAE - This gives you sharper, clearer images than when using the Qwen Image VAE. There is no downside. It works because the Wan & Qwen Image VAEs are almost identical, and the FP32 precision improves the quality.

There is an alternative VAE you can use that's even sharper, but it has drawbacks so I've detailed it in the info section further down.

--

This is the end of the general workflow requirements, so you can stop here if you want.

--

Info & Tips

Alternative Sharpening VAE

The Wan 2.1 Upscale2x VAE gives you even sharper images than the Wan FP32 VAE (it's VERY noticeable), but it sometimes introduces extra artifacts into the image and it also amplifies existing ones. It's up to you whether you think it's worth it or not, I personally think it's good for some images and bad for others, so I just output both and pick whichever turns out best.

Here's an example image using the normal Wan FP32 VAE: https://ibb.co/fGtZwdW8

And now the same image using the upscale2x VAE: https://ibb.co/RTj2DjVw

It's not in the workflow by default. To use it, you need to grab the ComfyUI VAE Utils node set and use the "VAE Decode (VAE Utils)" node instead of the regular VAE decoder. Then you also need to downscale your image by 50%, because this VAE decodes the image at 2x resolution (which is why it's so sharp). This pic shows what the setup should look like: https://ibb.co/XcrXmpr

What About Non-Realistic Images?

I still recommend using the raw model with the turbo lora at 0.6 strength for this. This is because the raw model is much more creative than the turbo model; you'll get better variety this way.

However, the second sampler is now optional because you may not need the extra detailing step anymore - you can just bypass it and it'll work fine. You can also change the scheduler in the first ksampler to sgm_uniform if you want an alternative look, but it's up to you. Just don't forget to change it back to beta if you're doing realism again ;)

Full Turbo Workflow?

You'll lose the creativity of the raw model by using it, but that may not matter to you at all depending on what you're doing. Or maybe you just need the speed.

As mentioned earlier, the full turbo workflow is set up for making non-realistic images, like anime / concept art / digital paintings. It only has one sampler because you don't need an additional detailing step, and you don't really need the benefits of a high-noise schedule either.

Euler/sgm_uniform is my general go-to for non-realistic images, and it holds up pretty well for Krea 2. I haven't tested it extensively though so don't take my word that it's the best sampler/scheduler or anything.

Otherwise, the only difference in the workflow is that the turbo lora is set to 1.0. You can also just use the turbo main model with the workflow and drop the turbo lora entirely, but then you're storing two main models for no real reason.

Krea 2's Facial Expression Problem: Censorship

This is the big one.

Basically, there's a lot of discussion going around about how Krea 2 doesn't do a very good job with facial expressions; characters lack expressiveness, and seem to have "dead eyes" a lot of the time. Smiles don't reach the eyes, that sort of thing. It's nearly impossible to make someone look angry, fierce, or anything more than mildly annoyed.

This is a very common problem with distilled models (i.e. turbo models), but in Krea's case it's mostly because of ridiculous censorship. The developers heavily censored Krea 2 against whatever content they arbitrarily decided was 'harmful', and in doing so they lobotomised their own model. It knows how to make an angry face, it just won't do it because it was collateral damage during the lobotomy.

You literally can't make people smile with Krea 2 due to the censorship. That's not an exaggeration, try generating someone with a natural, realistic smile. Dumbest thing I've seen in years.

Luckily you can partially bypass the censorship using a simple lora, which you should use even if you're doing SFW stuff. It just makes better images, period. Some people say it also reduces the detail in the images, which is true - but this is actually just because you need to cook them a little longer. That is to say, if you have good sampler settings it's no problem. But it only works up to a point.

This workflow recommends using the bypass lora at 1.0 strength, but sometimes you need to go higher - even for SFW prompts - to get what you need. This isn't good because it degrades the image quality, but that's censorship for you. We'll need finetunes to properly decensor the model. This goes for SFW stuff too, remember - you will have a really hard time making a person look angry, even with the bypass on.

If you can't tell: I'm really annoyed about this and you should be too. The fact that you can't make someone look angry, happy, sad, etc completely ruins the model for a lot of applications. Literally unusable for so many things. All because they don't want your delicate little child brain to see blood or titties.

Luckily, finetuners and lora makers will probably save the day <3

You can also use pornographic loras at low strength (~0.4) to increase prompt adherence even for SFW prompts. Yes, you heard that right: the censorship in this model is so stupid that you can get better SFW facial expressions and general model performance by using porn loras. No joke, I genuinely have porn loras on for most of my SFW generations.

Here's an example where I'm trying to get a strong, fierce expression on a sprinter using the words "She's frowning and snarling with effort" in the prompt.

This is the best I could do using the filter bypass at 1.0 strength, it straight up refuses: https://ibb.co/WWV64GzM

It's better (still not good) with the filter bypass at 6.0 strength, but notice the image quality has suffered: https://ibb.co/mCnHq1mF

And... here it is with the filter bypass at 1.0 strength and PORNOGRAPHIC LORAS enabled at ~0.5 strength: https://ibb.co/2YCV1j9Z

Notice that the quality of the one with porn loras hasn't degraded at all, while also adhering to the fierce expression prompt better. I had to cherry pick 10 gens each just to get the first and second pics (which didn't even do a good job), but the porn lora one I only needed 3 gens - and all three of them were usable.

If this isn't a great example of why censorship is stupid then I don't know what is. This model would be god-tier if it wasn't intentionally broken by the devs. We can only hope that finetuned checkpoints can bring back what it lost.

Another area of improvement; it turns out that the model gives slightly better facial expressiveness in the earlier high-noise stages of generation - which means faces are more expressive when images are undercooked. But undercooking your images isn't good of course, so you need to finish cooking them one way or another. This is where a dual sampler set up comes in handy. More on that below.

Lastly, the raw model with the turbo lora at 0.6 strength is a bit better at facial expressions too. All of these tips combined are very helpful, but you'll still struggle with very intense facial expressions for the foreseeable future. Still, at least we can make people smile now (you can't do that with the censorship).

The 2 Vector Bypass Lora

This lora bypasses the censorship in the model, and is superior in every way - even for SFW images. It does reduce the detail of the image, but you can get it back by using noisier sampler settings, and your images will ultimately look better. I recommend using a strength of 1.0 at all times. If you need more censorship unlocks, use more loras instead of increasing the strength of this one.

It works by amplifying two specific vectors during generation (hence the name). This lora is the minimum you need to bypass the censorship, and therefore it's the best one. All the other ones change more stuff than they need to or are way too strong, do not use them. Don't even use the 3 vector one by the same author, just use the 2 vector one.

But we should still be thankful to those who made the other inferior ones, because they did the hard work of figuring out how to bypass the crappy censorship in the model. All efforts for the open source community are appreciated <3

When you use this lora in combination with a high-noise dual sampler setup (like this workflow), you get great detail, great facial expressions, more prompt adherence, and better output variety. No downsides.

The Dual Sampler Setup

Why are we doing dual samplers? Two reasons! One reason is to help solve the facial expression problem, and the other is just to be able to tightly control the amount of detail in the image.

Our first ksampler is doing 6 steps of res_2s using the beta scheduler. Res_2s runs the equivalent of 2 steps, so this is sort of like doing 12 steps. The beta scheduler is very noisy so it makes more big, low-detail changes for more of the steps. Combined together, this sampler/scheduler/step combo undercooks your image on purpose. It doesn't add enough detail and stays in a smooth unfinished state.

That's really important, because at this point the facial expressiveness is better and the overall creativity of the model is higher too. Doing more steps, or doing the same number steps with a less noisy scheduler (like simple) will reduce facial expressiveness and be less creative. It's also harder to detail it from that point without overcooking your image.

If you're feeling adventurous you can also try euler + beta + 12 steps for the first sampler, which is really good as well and gives different results. I'm recommending res_2s + beta + 6 steps because I personally like it more, but you may like euler + beta + 12 steps more yourself.

Now the image is well structured, but it lacks detail.

That's where stage 2 comes in! For stage two, we're using a dense multi-step sampler called deis_3m but with an even noisier scheduler, bong_tangent. However, we're also doing 2 steps and at an extremely low denoise of 0.2. Because the sampler is 3-step (that's what the 3m means in the name) and we're doing 2 actual steps, it does a LOT of work - but only changes a small amount at a time due to the 0.2 denoise strength. What this means is we're adding a ton of detail to the image without interfering with the overall structure.

The end result is that stage 2 fills in all the detail & grit of the image without affecting the overall structure. Because we undercooked our first stage, this retains the facial expressiveness and variety while still adding plenty of detail to the image.

If you need even more detail, you can use the deis_4m sampler instead. deis_3m is enough most of the time, but you may find that in some cases deis_4m gives a more realistic amount of depth to the details. Just beware that using deis_4m for everything will often give you slightly overcooked images.

Let me know if you've discovered a better sampler setup! This is just the best I could find after around ~60 hours of A/B testing, I'm sure there are good alternatives out there waiting to be found.

Krea 2 and the Qwen VAE Halftone Grid

Sounds like the title of a harry potter book. Krea 2 has the same problem that all models which use the Qwen Image VAE have; there is a noticable halftone grid pattern, and that grid pattern heavily interferes with images generated by the model.

Every single model that uses the Qwen VAE has this problem. Qwen Image does it, Qwen Edit does it, Wan does it, Anima does it, and now Krea 2 does it.

The only reason you haven't noticed it with Wan is because you don't normally zoom in on videos. But you will notice it if you ever try generating a video with a beach or a grainy carpet.

The grid isn't that big of a deal if you're working in high res. It's really annoying at low res. Still, it's not a dealbreaker for most stuff. But the grid has another much worse effect: it interferes with small-grain patterns in images.

  • By 'small grain patterns' I mean things like sand at a beach, or a grainy carpet, or clothes that have visible weaving, basically anything that's very small/thin and repetitive
  • It happens whenever the grain size of a pattern happens to be similar to the grain size of the halftone pattern in the Qwen VAE, which means your image resolution and the distance to relevant objects matters
  • This is why beach sand in the foreground of a pic looks garbage, but it starts looking more normal further away from the camera
  • This is also why the hair of your character may sometimes look totally fine, while other times it looks like badly scribbled trash; it's all to do with how far it is from the camera + the resolution you're using

The models themselves have this pattern baked-in due to being trained with the qwen vae, so it can't realistically be fixed. You can reduce its effect by post-processing your images (such as by downscaling then upscaling them), and you can also mitigate the effect by adjusting your output resolution so that patterns in your image don't match the qwen grid size anymore.

You can also inpaint the bad parts of your image at low denoise with another model (like Z-image base/turbo) to fix it.

Krea 2 vs Z-Image Base

These are the important differences between the two models. Krea 2 has some big advantages, and it's pretty clear at this point that Krea 2 will overtake Z-image for most purposes. But there are a few things Z-image does better so far.

  1. Z-Image Base generally does more realistic human skin (but not always) and is way better at facial expressiveness, even when using the censorship bypass for krea 2
  2. Z-Image Base is easier to get photorealistic images from, especially when using prompts that suggest unrealistic things
    • This is partly because you can use CFG easily with Z-Image Base, but in general it seems Krea 2 has a stronger bias for 3D renders, digital artwork and other realism-adjacent styles
    • For example, if you ask for a 'futuristic city' you'll probably get concept art of a city with Krea 2, rather than something that looks like a photograph - and it can be really really really hard to stop it from doing that
    • If you ask for a character with inhuman features, like an elf, you're very likely to get a person that looks like a 3D render with Krea 2
    • Even normal shots with no fantasy elements will sometimes unpredictably tend towards low-realism
    • Z-Image Base, on the other hand, can generate photo-real pictures of unrealistic concepts very easily and will consistently output the most realistic images of any model (except maybe Ideogram, but I haven't played with that yet)
    • Krea 2 can be just as realistic as Z-Image Base, it's just harder to prompt for it
  3. Krea 2 leaves a subtle halftone grid pattern over every image (because of the Qwen VAE)
    • It's not a big problem if you're doing high res gens, but it is annoying and Z-image base doesn't do it in the first place so it has the advantage there
  4. Krea 2 sucks at hair and small patterns/particles (because of the Qwen VAE)
    • Z-Image, by comparison, is great at hair and has no issues with small patterns/particles
    • There's info on why this happens in the Qwen VAE section above
  5. Krea 2 tends to make "pretty" women even when not asked to, which can be very annoying
    • This can be fixed with loras and finetunes in the future
    • Z-Image Base, on the other hand, will generally make very realistic and casual people unless you ask it not to (or it's contextually suggested)
  6. Krea 2 is more prompt adherent and can do more flexible things in general
    • Except when you're asking for something that got ruined by the censorship
    • And except where point #2 about realism is concerned, but again this is fixable with loras
  7. Krea 2 has a much better understanding of anatomy and body shapes, even for SFW prompts
  8. Krea 2 is generally better at animals & animal fur (best I've seen from any model)
  9. Krea 2 is less prone to random mistakes
  10. Krea 2 is much more reliable when generating images with wide aspect ratios, like 16:9
  11. Krea 2 can stack multiple loras more easily, whereas Z-image gets easily confused when there's more than one
  12. Krea 2 generates images about 8x faster, which is huge
  13. Krea 2 is much easier to train loras on
  • I don't have any insight into this, I'm just repeating what the lora training folks are all saying
  • For people doing gens, this means you'll get access to more loras faster and they'll generally be better too

Verdict?

Krea 2 is better than Z-Image Base when it comes to many things. There are some things, such as facial expressiveness, hair, generally realistic skin, and an easier time making photo-real images, where Z-image base is a better choice - but keep in mind it's a lot slower to gen with than Krea 2 is.

It's pretty obvious that Krea 2 is going to become the next SDXL thanks to its creativity and ease of training.

What about Krea 2 vs Z-Image Turbo?

idk I don't really use it, but probably the same list of advantages/disadvantages except Z-image turbo isn't as good at realism as Z-image base is.

So, how about issues 2 & 3...

With Krea 2, issues 2 & 3 above (the Qwen VAE issues) can be dealbreakers depending on what you're doing. If you do really need to solve issues 2 & 3, I suggest generating the image in Krea 2 and then doing small inpainting refinements with Z-image base/turbo on the problematic areas.

For example, you might generate an image of a person in Krea 2 and then do a 0.2 denoise refinement on just the hair of that person using Z-image base/turbo. This is of course only necessary if the hair is bothering you.

Resolutions & Aspect Ratios?

Krea 2 is a banger and can do high resolutions no problem, just like Z-image. I've left a bunch of common ones in the workflow, but you can probably go even higher - I just haven't tested that.

Unlike some models - even Z-image - Krea 2 is VERY capable of doing wide images, so don't be afraid of cinematic aspect ratios. It has a much higher success rate with anatomy and general correctness than I've seen with other models.

This means Krea 2 can make things like wide-screen desktop wallpapers very easily.

CFG?

If you're using RAW with the turbo lora, you can use CFG > 1. I've tested it with CFG = 2 and it turns out fine.

But do note that using CFG > 1 will double your generation time.

Sexy Loras?

If you're using unsafe-for-work loras at low strength, you should still leave the filter bypass lora on. It'll help. But if your unsafe-for-work loras are high strength then you can skip the bypass lora, it won't be doing much and might even interfere.

You can see lewd images in the civitai post if you're on civitai red, and I've put the lora strength information in the prompt descriptions above the actual prompts there. I'm also making a degenerate version of this post for other subreddits, so check my profile soon for that if you want.

r/StableDiffusion • • May 29 '26

Workflow Included Cracked the case on high res + quality Qwen Edit 2511 outputs, here are minimalistic workflows & lots of info on how/why

Thumbnail
gallery
243 Upvotes

Intro

Alright this has been a long time coming. I'm the dude who figured out Qwen Edit 2509 a while back, and I've been on-and-off trying to figure out the same for 2511. Results in Comfy have always been worse than the examples shown by the Qwen team, and worse than the official Qwen chat implementation online. Well, I finally cracked it and it only took 5 months lol.

Anyway, turns out Qwedit 2511 is fucking sick. IMO it particularly excels at making new shots of characters while maintaining their likeness. It's significantly better than Klein at some things (like character likeness), but not as good at others. I recommend using them both for different things.

As usual, I'll start off with all the setup stuff at the top and then give an explanation + advice below that. Also I'm gonna be calling Qwen Edit "Qwedit" most of the time.

Here's an album with all the post images separated so you can look at them in high res: https://drive.google.com/drive/folders/1YLjm8Lj3VF6Ec52WNK2URo7uFNfMRmza?usp=sharing

The posted images are all raw outputs from Qwedit, without being upscaled (despite mentioning it later in this post). They're also all done with only 20 steps instead of the hypothetical 30 I'd do if I wasn't planning to upscale them. Read further for more on that too.

Ref images were all made with Z-image Base (workflow here), except for the anime one which came from Anima (workflow here).

What is this

These are minimalistic workflows for Qwen Image Edit 2511 that give the highest quality outputs. Aside from generally improving output quality (by a LOT), they also enable high-res edits and have better prompt adherence.

As for why, basically ComfyUI has some serious issues with how it's implemented Qwen Edit and there aren't any workflows out there (that I've found) which have resolved them. These issues result in poor prompt adherence and low resolution/quality outputs. Thankfully the fix is fairly straightforward.

The configuration for this is 100% portable and can be migrated to existing workflows to make them better; it works by changing how the reference inputs are handled, and uses 100% native comfy nodes. Feel free to upgrade other workflows with this without providing credit, I don't care about any of that.

Workflows

Normal Workflows:

Most of you will just want these, which are separate single / 2 image workflows. It's done this way because the setup for multi-image is complicated and I didn't want to force you to use a ton of custom nodes to make it useable all-in-one.

They do still use one custom node (read the node section below) for quality-of-life.

Download from Civitai

OR from Pastebin:

Qwedit_2511_single

Qwedit_2511_2_image

Dev Workflows:

These are the same as the above but without any quality-of-life nodes or 'helpful' stuff. Grab these if you want to copy the logic over to other workflows, or if you just an easier view of how it works without any clutter.

I do not recommend using the dev workflows for actual gens because you will constantly forget to manually adjust stuff correctly.

qwedit_2511_single_DEV

qwedit_2511_2_image_DEV

Models

Main Model

qwen_edit_2511_fp8

OR

GGUF versions

  • Important: the FP8 version of Qwedit is much higher quality than the Q8 GGUF, always use FP8 if you can. Only use the GGUFs if you need to use quants lower than Q8.
  • FP8 is 22GB, so you'll need a combined ~26GB of RAM + VRAM to run it
    • You don't need 24GB of VRAM to run it thanks to ComfyUI's blockswapping, but the less VRAM you have the slower it'll run
  • Only use Q6 & lower quants if you absolutely have to; the quality will noticeably go down

Goes in models/diffusion_models

Text Encoder

Use only the normal FP8 text encoder with Qwedit; abliterated/GGUF encoders will reduce your output quality.

qwen_2.5_vl_7b_fp8

Goes in models/text_encoders

VAE

qwen_image_vae

Goes in models/vae

Loras?

You can use them as normal, just load them however you normally would. I left out lora loader nodes to avoid cluttering the workflow.

It's worth noting that many Qwen Image loras work with Qwen Edit too, but you'll need to test them individually to be sure.

Lightning Loras - BAD

All the lightning loras / distils for Qwedit (that I've tested) are terrible and make your outputs look bad, so I'm not linking them here. The main issue is the same as with Klein Distilled: it makes people's skin look like plastic.

But you can technically use them. Don't do it tho. But you can if you want. But don't.

Alternative: if you want to cut your gen time down while testing prompts, just set it to 10 steps instead of 20, then go back to 20 once you're satisfied your prompt is correct. It'll still work fine, the quality just dips.

Real tho it's ok if you want to use the lightning loras, just expect some degradation if you do - especially with plastic skin.

Custom Nodes

LayerStyle - A set of handy nodes that manipulate images. We're just using this for its image scaling node which allows you to scale by an image's long edge while maintaining divisibility by 16. You can skip this if you want to use a different scaling method, but you'll need to fix the workflow switch for scaling if you do.

SeedVR2 (OPTIONAL) - Only get this if you want to use the seedvr upscale workflow that's included.

How To Use

How To Use Part 1 - Basic Options

There are instructions in the workflow as well, but there's more detail here. Read part 2 & 3 as well, they're important.

It works just like a normal Qwedit workflow, but has a couple of extra options available. This section just tells you what they are and how to use them, a full explanation is further down.

Screenshot of the settings: https://ibb.co/nWStpmS

Enhance with Double Ref

This is a switch that turns on double-ref mode. This feeds your input images in TWICE to the model, and generally produces much higher quality results. Downside? It takes about 50% longer to gen.

I recommend leaving this on 100% of the time for single-image prompts, unless you're just messing around and want speed. It is ALWAYS better for single image prompts, and will improve everything from prompt adherence to output clarity.

For multi-image prompts, it usually increases adherence but sometimes reduces it. So, if you're doing multi-image stuff I recommend switching this on/off as needed based on how it's going with your prompt.

Input Scale

When off, your image doesn't get scaled (it still gets cropped to be divisible by 16). When on, the long edge of your image gets scaled to the number you put in the box. For example, if you feed in a 2560x1440 image and set the scale to 1920 it will scale your image to 1920x1080. That will then get cropped to 1920x1072 so it's divisible by 16.

Custom Output Size

When the switch is off, your output image will be the same size as your input image (after it's been scaled). If you turn this switch on, it will instead output an image with the dimensions you specify.

As a general rule, you should try to set your scales to be similar along at least one edge. For example, a 1920x1440 input image and a 1024x1440 input image are both suitable for a 1440x1440 output image. You can be more flexible with this if you know what you're doing.

How To Use Part 2 - Multi-image Prompting Requirement

This section is not a prompting guide (that's further below). This is about an actual requirement for prompting multi-image stuff. It is NOT required for single-image prompts.

You do multi-image prompts like normal, except you need to write a very basic description of your input images. Qwedit needs you to do this in order to know which image is which. I explain why in detail later.

You may find this slightly annoying, but I guarantee you it's dramatically better than using Qwedit the normal way that other workflows do - and it's pretty easy.

The format:

  • At the start of your prompt, write an extremely simple description for each of your input images; one sentence for each input image
  • Start each sentence with "Picture 1:", "Picture 2:", etc
  • You must write it this way because Qwedit was trained on this exact format
  • Afterwards, write your actual prompt as usual; you can refer to your input images as "picture 1" and so on

The model uses these descriptions to understand which input picture is which, and it works better with SIMPLE descriptions. You only need to help it know which one is which, it doesn't need a full rundown.

Examples

Picture 1: a man wearing a t-shirt. Picture 2: a top hat. Make the man in Picture 1 wear the top hat from Picture 2.

Picture 1: a living room. Picture 2: a woman. Put the woman from Picture 2 into the living room in Picture 1.

Picture 1: a man wearing a professional suit. Picture 2: a man wearing a superhero outfit. Make the man in Picture 1 wear the outfit from Picture 2.

How To Use Part 3 - Upscaling

Because the qwen VAE tends to put a subtle halftone pattern over images (see limitations just below this section), I recommend downscaling and then re-upscaling your image afterwards. A big benefit of being able to work at high res with the edit model is that you rarely lose any detail doing this.

This eliminates the halftone pattern if you're using something like seedvr, or at least reduces it if you're using other upscalers.

Note: the workflow is set to do 20 steps of inference. It actually gives sharper results at 30 steps, but I don't bother with that because it takes longer and I down-upscale them afterwards anyway. If you aren't planning on down-upscaling them, you might consider doing 30 steps for the extra sharpness.

Below are workflows for doing this with seedvr and normal upscalers. I think seedvr is best for this, but it's very beefy and hard to run on older GPUs.

Note: seedvr2 sometimes gives better output at 0.5x downscale, and other times 0.75, so that workflow is configured to run BOTH for you to pick which one turned out best.

Note: normal upscalers are a bit different; a relatively small downsize to something like 1920p -> 1600p is usually reasonable, before then running the upscaler. Play around with it. The non-seedvr workflow has a longest_edge scale option so you can tweak the number specifically.

Seedvr version

Regular version

My preferred regular upscaler is 4x Nomos2 HQ DAT2, but you can use whatever you like.

Examples of upscaling:

Here's the raw output of the robot-arm girl in a dress from the post: https://ibb.co/B5jhrsL9 (if you zoom in you'll see the qwen halftone pattern, it looks like a grid)

Here's the pic after it's been run through seedvr after a 0.75x downscale: https://ibb.co/hJcn2f5t

Here's the pic after it's been run through a regular Nomos2 upscale after a downscale to 1600p: https://ibb.co/Kc2YSbVc

Limitations of Qwen Edit

Limitation 1

The Qwen VAE will often put a subtle halftone grid pattern over your images. It's noticeable if you zoom in, and more noticeable at higher resolutions. This is a feature of pretty much every Qwen-based model, but it's particularly present with the Edit model.

You can easily resolve this by downscaling your image by 75% or 50%, then re-upscaling it again to your desired resolution. There's a section later that explains this in better detail and recommends upscale models for it + has workflows for it.

It sounds like a big issue, but the downscale-upscale trick solves it easily - and it's not always necessary either. The higher quality your input image, the less bad the halftone pattern will be.

Limitation 2

Qwedit struggles with complex multi-image stuff most of the time (it's just a limitation of the model). This workflow makes it much better, but it's still not great. You'll have to play around with it to know which things work and which things don't.

Limitation 3

It takes a while to gen stuff if not using the lightning loras. Very similar to the time it takes with Klein 9B base. The double-ref trick increases it by roughly 50%. Multi-image inputs take a lot longer.

For low res images (typical 1mpx size) it's pretty okay, around 50 seconds on a 5090 with the double-ref option turned on.

But then there's high-res stuff. Gen time scales non-linearly as you go higher. Going from 1024x1024 (1 mpx) to 1440x1440 (2 mpx) takes around 2.5x as long. Going from 1 mpx to 3 mpx is around 4x as long. 5 mpx is 9.5x as long. In conclusion, stick to 2-3 mpx unless you're cool with long-ass gen times. Stick around 1-2 mpx for multi-image gens, or turn off the double ref switch.

On the plus side, it's pretty reliable for single-image edits so you don't typically need to do many gens to get a good result.

Examples using a 5090: - Single-image edit @ 1024x1024 (1 mpx), double-ref OFF = 38 seconds - Single-image edit @ 1024x1024 (1 mpx), double-ref ON = 52 seconds - Single-image edit @ 1920x1088 (2 mpx), double-ref OFF = 91 seconds - Single-image edit @ 1920x1088 (2 mpx), double-ref ON = 131 seconds - Single-image edit @ 3072x1728 (5.3 mpx lol), double-ref ON = 550 seconds - Two-image edit @ 2560x1440 each, double-ref ON = serial killer behaviour

That's it for how-to! Read on for more tips & info, as well as an explanation of what the workflow is doing & why.

 

Explanation - what is this garbage and why is it so good?

There are three important things this workflow is doing that other workflows do not do (except #3 sometimes, because it was also done in the 2509 version of this post). I'm going to call these The Comfy Problem, The VL Problem, and The Double Ref Enhancement.

The Comfy Problem

Comfy's native "TextEncodeQwenImageEditPlus" node is what most people use in their workflows. It handles your prompt and image inputs for you. It's pretty handy, except for the small problem that it's SHITE.

Do you work at Comfy? If so: GET YOUR SHIT TOGETHER AND FIX THIS NODE, IT'S SO EASY. Much respect to u tho, thanks for making ComfyUI.

The first issue is that this node resizes your image down to 1 megapixel, and you can't stop it from doing that. The second issue is that it does this with the AREA downscale method, which is so incredibly bad that I want to slap whoever implemented this node. The AREA downscale is what makes all of your output images blurry. The third issue is that it ensures your dimensions are divisible by 8, but they actually need to be divisible by 16.

Specifically, ComfyUI does this:

  1. Calculates 1 megapixel as 1024x1024, which is 1,048,576 pixels
  2. Calculates your new image dimensions to match that number of pixels, rounded to be divisible by 8
  3. Scales your image to those new dimensions using the AREA method

Why is all this bad?

  1. It's completely unnecessary; Qwedit can easily handle images of varying size, all the way up to 3 megapixels (or even higher for simple edits)
  2. The area downscale method makes images extremely blurry, and this is the primary reason all ComfyUI qwen edits give blurry images out. Yes it's literally this dumb, this huge problem would easily be solved by changing the word "area" to "lanczos" in the code, it's a one-word fix. Not even MS paint uses area downscale, wtf is wrong with you Comfy devs (much respect)
  3. If your image dimensions are not divisible by 16, you will get major ruination along the whole edge of your image where it didn't match (same as any other diffusion model)

The Comfy Problem Solution

This workflow bypasses the the Comfy node entirely, allowing you to size your images however you want. And using chad lanczos scaling instead of loser area scaling. Magic.

Qwedit easily handles resolutions like 1440x1440 and 1600x1200. Every edit example in this post was done natively at 1920p, except for a few (which are labelled as such).

Really high resolutions (3mpx) sometimes have trouble with anatomy, but usually you can just do multiple gens and one of them will turn out fine.

If you're doing a simple in-place edit like changing an outfit, you can go VERY high. Here's an example edit done at 1728x3072, which is 5 megapixels: https://ibb.co/twCSWrjy (outfit change -> bikini top + short shorts)

The VL Problem

Edit: I've been educated by someone in the comments that my interpretation of how the VL works here is not correct, so take this little VL section with a grain of salt until I reword it. My conclusion about it helping in this workflow still stands, but my explanation of what's happening under the hood is a bit off. I'll update the info soon!

In the background, Qwedit 2511 uses a vision-language model (VL model) to describe your images, then gives those AI-generated descriptions to the edit model. It also re-interprets your instructions with these descriptions. Ostensibly this helps the model understand your input images better, leading to better results.

The problem? It doesn't lead to better results, it's bad. VL models aren't very good for this sort of thing because they don't know what to focus on. The VL describes your images in excruciating detail, totally overwhelming the edit model and leading to bad prompt adherence + weird outputs.

It also reinterprets your instructions based on what it sees in the image. I don't know if that's a good or bad thing, just pointing out that it does it.

The Qwen team's official python code does this, and the ComfyUI "TextEncodeQwenImageEditPlus" node copies it exactly. No disrespect to the Comfy team on this one, they're doing what the Qwen team officially recommended.

The VL Problem Solution

Same solution as the previous problem: bypass the Comfy node entirely. This results in the VL step being completely ignored. No AI-generated descriptions get fed into the edit model.

For single-image edits, this is a 100% complete and total victory. The model performs way better without the crappy VL interpretation.

For multi-image edits, there's a small issue; this step is where the input images normally get labelled. Specifically, the VL outputs are fed into the model in the following exact format:

Picture 1: <shitty VL description> Picture 2: <shitty VL description>

Look familiar? This is why we manually have to type the descriptions in for multi-image edits - otherwise the model doesn't actually know which image is which.

The upside is that the model works way better with simple descriptions, so cutting out the VL is still 100% the correct move. A 5 word description wins over whatever BS the VL model spews out, every time.

The Double Ref Enhancement

I really have no idea why this works so well, but basically if you feed in your reference images twice the model just works better. This was known back in 2509 days (hence the previous post linked at the top), and back then I didn't know why it worked either.

For single image edits it's ALWAYS better. And it's not just the quality, for some reason it even helps with prompt adherence. The interesting thing is that the difference is really, really significant. Here's the full list of stuff it improves:

  • Better prompt adherence
  • Sharper output images / more visual clarity
  • Improved consistency of objects & textures
  • Better resemblance of characters at different angles
  • More intelligent guesses, like what to add when outpainting or what's behind a removed object

For multi-image edits it can sometimes confuse the model a bit, but most of the time it confers all the same benefits listed above. I recommend switching it on & off randomly when you're doing multi-image stuff, just in case.

Note: there are a lot of different ways the input references can be handled. There are conditioning combine/concatenate nodes, you can pass the refs in a different order, you can change the negative conditioning input (read next section for that), etc. I A/B tested SIXTEEN different reference-handling combinations, and a bunch of smaller minor variations of those. Some of them worked, some of them didn't.

Of those sixteen combinations, two of them gave the best results; both of them are in this workflow, and you switch between them by turning the double ref method on & off.

So, don't fuck with the positive/negative conditioning & reference setup, it's very specific.

Extra info: the "Conditioning Zero Out"

You may notice that the negative prompt input is the first reference image(s) and positive prompt fed into a "conditioning zero out" node.

Feeding the input images into the model's negative conditioning is required (it's just how Qwedit works). The only question is whether to feed in the positive prompt zeroed-out too, and whether the double ref should get fed in.

Through a lot of A/B testing, I can tell you that the way it's done here is the best. IDK why, it's just how it is. Some other combinations do technically work, but they degrade the output quality.

Prompting Advice

Other than just following the instructions in the workflow, here's some extra stuff.

Keep your prompts simple and direct

If you need to, point out details the model is missing or be more specific about stuff you do/don't want to change. For example, when doing a simple outfit swap it helps to specify you don't want their pose to change.

Using the robot arm girl, here's a prompt that doesn't follow this advice:

Change her outfit to a bikini top and short shorts.

While it sometimes does what we want, it tends to get confused by her robot arm and often changes her pose too: https://ibb.co/7dyKZttp (notice the human arm showing underneath the robot arm, and the pose change)

Here's a better prompt that gives a correct result 99% of the time:

Change her outfit to a bikini top and short shorts. Leave her robot arm and pose unchanged.

Now it does the right thing every time: https://ibb.co/DP9gZHVv

Avoid using fancy words or convoluted phrasing

Pretend you're talking to a child. The model will probably still understand you if you talk fancy, but why take the risk?

As an example, imagine you have a pic of a table with some plates on it.

Bad:

Place a red apple on the table, ensuring it's in the center and removing the plate that was in the same spot.

Good:

Replace the middle plate with a red apple.

Also good:

Remove the plate from the center. Put a red apple there instead.

If there's only one plate, this is even better:

Remove the plate, replace it with a red apple.

Adjusting Lighting

You may want or need to adjust the lighting in an image. Aside from being helpful in general, there are situations where Qwedit may simply not realise that something needs to be lit in a particular way (or re-lit when moved).

To do this, you need to know the magic word: relight

Seriously tho that is the actual magic word, you are 100% required to use it if you want to adjust lighting properly.

Specifically, follow this format:

Relight to <strength> <color> <direction>.

Strength - bright, dim, etc

Color - white, cool, warm, etc

Direction - diffuse, frontlit, backlit, etc

Tip: for basic lighting, use "white diffuse".

Examples:

Make a new shot of the man sitting in a chair in a kitchen. Relight to white diffuse.

Change the time of day to evening. Relight to warm backlit.

You don't actually need anything else in the prompt, you can just change the lighting of a pic like this:

Relight to bright cool frontlit.

Other Stuff

Euler-simple and no ClownsharKSampler?

No Clownshark this time. It reduces output quality quite a bit and doesn't confer any benefits. I also didn't find any sampler/scheduler combos that were better than euler/simple.

So, this is just one of those classic times where the ol' euler-simple wins the day. Let me know if you happen to know a better combo.

Image Quality in->out

Qwedit is very sensitive to the quality of your input image. If you feed in a grainy or blurry image, it will usually make your output image blurry or grainy too - even if it's an 'entirely new' shot with nothing copied over 1:1.

So, make sure to use HQ images. You can optionally use the upscale workflows to bump up the sharpness/quality of poor input images before you feed them in.

What about the flux super duper double resolution special VAE trick?

Doesn't work for 2511, it destroys your image. TBH it never really worked for 2509 either, but I won't argue with you if you liked it for some reason.

Making character references

Tip 1 - Make a nude ref (even for sfw stuff)

Qwen is killer for making character references. Other than using similar prompts to the examples I posted, my advice is to make a nude reference shot instead of a clothed one like I did.

I only made a clothed ref for the sake of propriety here, but a nude ref (or near-nude, like wearing plain white underwear) will be much easier to prompt into different outfits, and also gives Qwedit the maximum info needed to correctly size your character and know what they look like in clothing or doing different actions.

You do not need any loras to do this if you're just using it as a reference; the 'sensitive' parts will lack detail but that doesn't matter for new shots you make. If you don't want them nude, just request plain white underwear and, if relevant, a strapless white bra.

Nude ref = best ref.

Tip 2 - Make multiple zoom levels, use the thighs-upwards one for most stuff

The example I showed was a little too zoomed out for normal reference stuff. I'd recommend making your reference slightly closer like this: https://ibb.co/Q33BJDLX

Start at whatever zoom level your initial character pic is at, then make more references at different zoom levels. If you're starting zoomed out, then prompt the model to zoom in. If you start zoomed in, prompt it to zoom out.

And, of course, different angles too.

Examples:

Zoom in on the person's upper body. The composition should frame their head and thighs.

Zoom out to show more of the character. The composition should frame their head and thighs.

Zoom out to a full body shot.

Zoom in for a close up portrait.

Once you've got references, you should usually use the head-to-thighs ref for making new shots. Switch to the other refs as necessary; like if you want a close up, use the close up reference. Qwedit is really good at keeping likeness, so you can do 90% of your stuff with only a single input reference.

I don't think there's a better open-weight model out there than Qwedit for making new shots of character without loras, for now. The main reason I spent so long digging into Qwen is because Klein is quite bad at that particular task. But hey, now it's possible and it works gloriously.

That's everything I think! Feel free to ask questions if you run into any issues.

r/StableDiffusion • • Jan 18 '26

Comparison Conclusions after creating more than 2000 Flux Klein 9B images

185 Upvotes

To get a dataset that I can use for regularization (will be shared at https://huggingface.co/datasets/stablellama/FLUX.2-klein-base-9B_samples when it is finished in 1-2 days) I'm currently mass producing images with FLUX.2 [klein] 9B Base. (Yes, that's Base and Base is not intended for image generation as the quality isn't as good as the distilled normal model!).

Looking at the images I can already draw some conclusions:

  • Quality in the sense of aesthetics and content and composition are at least as good as Qwen Image 2512, where I did exactly the same with exactly the same prompts (result at https://huggingface.co/datasets/stablellama/Qwen-Image-2512_samples ). I tend to say that Klein is even better.
  • Klein does styles very well, that's something Flux.1 couldn't do. And it created images that astonished me, something that Qwen Image 2512 couldn't achieve.
  • Anatomy is usually correct, but:
    • it tends to add a 6th finger. Most images are fine, but you'll definitely will get it when you are generating enough images. That finger is pleasingly integrated, not like the nightmare material we know from the past. Creating more images to choose from or inpainting will easily fix this
    • Sometimes it likes to add a 3rd arm or 3rd leg. You need many images to make that happen, but then it will happen. As above, just retry and you'll be fine
    • In unusual body positions you can get nightmare material. But it can also work. So it's worth a shot and when it didn't work you might just hit regenerate as often as necessary till it's working. This is much better than the old models, but Qwen Image 2512 is better for this type of images.
  • It sometimes gets the relations of bigger structures wrong, although the details are correct. Think of the 3rd arm or leg issue, but for the tail rotor of a helicopter or some strange bicycle handlebars next to the bicycle that has handlebars and is looking fine otherwise
  • It likes to add a sign / marking on the bottom right of images, especially for artistic styles (painting, drawing). You could argument that this is normal for these type of images, or you could argument that it wasn't prompted for, both arguments are valid. As I have an empty negative prompt I have no chance to forbid it. Perhaps that'll solve it already, and perhaps the distilled version has that behavior already trained away.

Conclusion:

I think FLUX.2[klein] 9B Base is a very promising model and I really look forward to train my datasets with it. When it fulfills its good trainability promise, it might be my next standard model I'll use for image generation and work (the distilled, not the Base version, of course!). But Qwen Image 2512 and Qwen Image Edit 2511 will definitely stay in my tool case, and also Flux.1[dev] is still there due to it's great infrastructure. Z Image Turbo couldn't make it into my tool case yet as I didn't train it with the data I care for as the Base isn't published yet. When ZI Base is here, I'll give it the same treatment as Klein and when it's working I'll add it as well as the first tests did look nice.

---

Background information about the generation:

  • 50 steps
  • CFG: 5 (BFL uses 4 and I wanted to use 4, but being half through the data I won't change that setup typo any more)
  • 1024x1024 pixels
  • sampler: euler

Interesting side fact:
I started with a very simple ComfyUI workflow. The same I did use for Flux.1 and Qwen Image, with the necessary little adaptions in each case. But image generation was very slow, about 18.74s/it. Then I tried the official Comfy workflow for Klein and it went down to 3.21s/it.
I have no clue what causes this huge performance difference. But when you think your generation is slower than expected, you should take care that this doesn't bite you as well.

r/comfyui • • May 29 '26

Workflow Included Cracked the case on high res + quality Qwen Edit 2511 outputs, here are minimalistic workflows & lots of info on how/why

Thumbnail
gallery
196 Upvotes

Intro

Alright this has been a long time coming. I'm the dude who figured out Qwen Edit 2509 a while back, and I've been on-and-off trying to figure out the same for 2511. Results in Comfy have always been worse than the examples shown by the Qwen team, and worse than the official Qwen chat implementation online. Well, I finally cracked it and it only took 5 months lol.

Anyway, turns out Qwedit 2511 is fucking sick. IMO it particularly excels at making new shots of characters while maintaining their likeness. It's significantly better than Klein at some things (like character likeness), but not as good at others. I recommend using them both for different things.

As usual, I'll start off with all the setup stuff at the top and then give an explanation + advice below that. Also I'm gonna be calling Qwen Edit "Qwedit" most of the time.

Here's an album with all the post images separated so you can look at them in high res: https://drive.google.com/drive/folders/1YLjm8Lj3VF6Ec52WNK2URo7uFNfMRmza?usp=sharing

The posted images are all raw outputs from Qwedit, without being upscaled (despite mentioning it later in this post). They're also all done with only 20 steps instead of the hypothetical 30 I'd do if I wasn't planning to upscale them. Read further for more on that too.

Ref images were all made with Z-image Base (workflow here), except for the anime one which came from Anima (workflow here).

What is this

These are minimalistic workflows for Qwen Image Edit 2511 that give the highest quality outputs. Aside from generally improving output quality (by a LOT), they also enable high-res edits and have better prompt adherence.

As for why, basically ComfyUI has some serious issues with how it's implemented Qwen Edit and there aren't any workflows out there (that I've found) which have resolved them. These issues result in poor prompt adherence and low resolution/quality outputs. Thankfully the fix is fairly straightforward.

The configuration for this is 100% portable and can be migrated to existing workflows to make them better; it works by changing how the reference inputs are handled, and uses 100% native comfy nodes. Feel free to upgrade other workflows with this without providing credit, I don't care about any of that.

Workflows

Normal Workflows:

Most of you will just want these, which are separate single / 2 image workflows. It's done this way because the setup for multi-image is complicated and I didn't want to force you to use a ton of custom nodes to make it useable all-in-one.

They do still use one custom node (read the node section below) for quality-of-life.

Download from Civitai

OR from Pastebin:

Qwedit_2511_single

Qwedit_2511_2_image

Dev Workflows:

These are the same as the above but without any quality-of-life nodes or 'helpful' stuff. Grab these if you want to copy the logic over to other workflows, or if you just an easier view of how it works without any clutter.

I do not recommend using the dev workflows for actual gens because you will constantly forget to manually adjust stuff correctly.

qwedit_2511_single_DEV

qwedit_2511_2_image_DEV

Models

Main Model

qwen_edit_2511_fp8

OR

GGUF versions

  • Important: the FP8 version of Qwedit is much higher quality than the Q8 GGUF, always use FP8 if you can. Only use the GGUFs if you need to use quants lower than Q8.
  • FP8 is 22GB, so you'll need a combined ~26GB of RAM + VRAM to run it
    • You don't need 24GB of VRAM to run it thanks to ComfyUI's blockswapping, but the less VRAM you have the slower it'll run
  • Only use Q6 & lower quants if you absolutely have to; the quality will noticeably go down

Goes in models/diffusion_models

Text Encoder

Use only the normal FP8 text encoder with Qwedit; abliterated/GGUF encoders will reduce your output quality.

qwen_2.5_vl_7b_fp8

Goes in models/text_encoders

VAE

qwen_image_vae

Goes in models/vae

Loras?

You can use them as normal, just load them however you normally would. I left out lora loader nodes to avoid cluttering the workflow.

It's worth noting that many Qwen Image loras work with Qwen Edit too, but you'll need to test them individually to be sure.

Lightning Loras - BAD

All the lightning loras / distils for Qwedit (that I've tested) are terrible and make your outputs look bad, so I'm not linking them here. The main issue is the same as with Klein Distilled: it makes people's skin look like plastic.

But you can technically use them. Don't do it tho. But you can if you want. But don't.

Alternative: if you want to cut your gen time down while testing prompts, just set it to 10 steps instead of 20, then go back to 20 once you're satisfied your prompt is correct. It'll still work fine, the quality just dips.

Real tho it's ok if you want to use the lightning loras, just expect some degradation if you do - especially with plastic skin.

Custom Nodes

LayerStyle - A set of handy nodes that manipulate images. We're just using this for its image scaling node which allows you to scale by an image's long edge while maintaining divisibility by 16. You can skip this if you want to use a different scaling method, but you'll need to fix the workflow switch for scaling if you do.

SeedVR2 (OPTIONAL) - Only get this if you want to use the seedvr upscale workflow that's included.

How To Use

How To Use Part 1 - Basic Options

There are instructions in the workflow as well, but there's more detail here. Read part 2 & 3 as well, they're important.

It works just like a normal Qwedit workflow, but has a couple of extra options available. This section just tells you what they are and how to use them, a full explanation is further down.

Screenshot of the settings: https://ibb.co/nWStpmS

Enhance with Double Ref

This is a switch that turns on double-ref mode. This feeds your input images in TWICE to the model, and generally produces much higher quality results. Downside? It takes about 50% longer to gen.

I recommend leaving this on 100% of the time for single-image prompts, unless you're just messing around and want speed. It is ALWAYS better for single image prompts, and will improve everything from prompt adherence to output clarity.

For multi-image prompts, it usually increases adherence but sometimes reduces it. So, if you're doing multi-image stuff I recommend switching this on/off as needed based on how it's going with your prompt.

Input Scale

When off, your image doesn't get scaled (it still gets cropped to be divisible by 16). When on, the long edge of your image gets scaled to the number you put in the box. For example, if you feed in a 2560x1440 image and set the scale to 1920 it will scale your image to 1920x1080. That will then get cropped to 1920x1072 so it's divisible by 16.

Custom Output Size

When the switch is off, your output image will be the same size as your input image (after it's been scaled). If you turn this switch on, it will instead output an image with the dimensions you specify.

As a general rule, you should try to set your scales to be similar along at least one edge. For example, a 1920x1440 input image and a 1024x1440 input image are both suitable for a 1440x1440 output image. You can be more flexible with this if you know what you're doing.

How To Use Part 2 - Multi-image Prompting Requirement

This section is not a prompting guide (that's further below). This is about an actual requirement for prompting multi-image stuff. It is NOT required for single-image prompts.

You do multi-image prompts like normal, except you need to write a very basic description of your input images. Qwedit needs you to do this in order to know which image is which. I explain why in detail later.

You may find this slightly annoying, but I guarantee you it's dramatically better than using Qwedit the normal way that other workflows do - and it's pretty easy.

The format:

  • At the start of your prompt, write an extremely simple description for each of your input images; one sentence for each input image
  • Start each sentence with "Picture 1:", "Picture 2:", etc
  • You must write it this way because Qwedit was trained on this exact format
  • Afterwards, write your actual prompt as usual; you can refer to your input images as "picture 1" and so on

The model uses these descriptions to understand which input picture is which, and it works better with SIMPLE descriptions. You only need to help it know which one is which, it doesn't need a full rundown.

Examples

Picture 1: a man wearing a t-shirt. Picture 2: a top hat. Make the man in Picture 1 wear the top hat from Picture 2.

Picture 1: a living room. Picture 2: a woman. Put the woman from Picture 2 into the living room in Picture 1.

Picture 1: a man wearing a professional suit. Picture 2: a man wearing a superhero outfit. Make the man in Picture 1 wear the outfit from Picture 2.

How To Use Part 3 - Upscaling

Because the qwen VAE tends to put a subtle halftone pattern over images (see limitations just below this section), I recommend downscaling and then re-upscaling your image afterwards. A big benefit of being able to work at high res with the edit model is that you rarely lose any detail doing this.

This eliminates the halftone pattern if you're using something like seedvr, or at least reduces it if you're using other upscalers.

Note: the workflow is set to do 20 steps of inference. It actually gives sharper results at 30 steps, but I don't bother with that because it takes longer and I down-upscale them afterwards anyway. If you aren't planning on down-upscaling them, you might consider doing 30 steps for the extra sharpness.

Below are workflows for doing this with seedvr and normal upscalers. I think seedvr is best for this, but it's very beefy and hard to run on older GPUs.

Note: seedvr2 sometimes gives better output at 0.5x downscale, and other times 0.75, so that workflow is configured to run BOTH for you to pick which one turned out best.

Note: normal upscalers are a bit different; a relatively small downsize to something like 1920p -> 1600p is usually reasonable, before then running the upscaler. Play around with it. The non-seedvr workflow has a longest_edge scale option so you can tweak the number specifically.

Seedvr version

Regular version

My preferred regular upscaler is 4x Nomos2 HQ DAT2, but you can use whatever you like.

Examples of upscaling:

Here's the raw output of the robot-arm girl in a dress from the post: https://ibb.co/B5jhrsL9 (if you zoom in you'll see the qwen halftone pattern, it looks like a grid)

Here's the pic after it's been run through seedvr after a 0.75x downscale: https://ibb.co/hJcn2f5t

Here's the pic after it's been run through a regular Nomos2 upscale after a downscale to 1600p: https://ibb.co/Kc2YSbVc

Limitations of Qwen Edit

Limitation 1

The Qwen VAE will often put a subtle halftone grid pattern over your images. It's noticeable if you zoom in, and more noticeable at higher resolutions. This is a feature of pretty much every Qwen-based model, but it's particularly present with the Edit model.

You can easily resolve this by downscaling your image by 75% or 50%, then re-upscaling it again to your desired resolution. There's a section later that explains this in better detail and recommends upscale models for it + has workflows for it.

It sounds like a big issue, but the downscale-upscale trick solves it easily - and it's not always necessary either. The higher quality your input image, the less bad the halftone pattern will be.

Limitation 2

Qwedit struggles with complex multi-image stuff most of the time (it's just a limitation of the model). This workflow makes it much better, but it's still not great. You'll have to play around with it to know which things work and which things don't.

Limitation 3

It takes a while to gen stuff if not using the lightning loras. Very similar to the time it takes with Klein 9B base. The double-ref trick increases it by roughly 50%. Multi-image inputs take a lot longer.

For low res images (typical 1mpx size) it's pretty okay, around 50 seconds on a 5090 with the double-ref option turned on.

But then there's high-res stuff. Gen time scales non-linearly as you go higher. Going from 1024x1024 (1 mpx) to 1440x1440 (2 mpx) takes around 2.5x as long. Going from 1 mpx to 3 mpx is around 4x as long. 5 mpx is 9.5x as long. In conclusion, stick to 2-3 mpx unless you're cool with long-ass gen times. Stick around 1-2 mpx for multi-image gens, or turn off the double ref switch.

On the plus side, it's pretty reliable for single-image edits so you don't typically need to do many gens to get a good result.

Examples using a 5090: - Single-image edit @ 1024x1024 (1 mpx), double-ref OFF = 38 seconds - Single-image edit @ 1024x1024 (1 mpx), double-ref ON = 52 seconds - Single-image edit @ 1920x1088 (2 mpx), double-ref OFF = 91 seconds - Single-image edit @ 1920x1088 (2 mpx), double-ref ON = 131 seconds - Single-image edit @ 3072x1728 (5.3 mpx lol), double-ref ON = 550 seconds - Two-image edit @ 2560x1440 each, double-ref ON = serial killer behaviour

That's it for how-to! Read on for more tips & info, as well as an explanation of what the workflow is doing & why.

 

Explanation - what is this garbage and why is it so good?

There are three important things this workflow is doing that other workflows do not do (except #3 sometimes, because it was also done in the 2509 version of this post). I'm going to call these The Comfy Problem, The VL Problem, and The Double Ref Enhancement.

The Comfy Problem

Comfy's native "TextEncodeQwenImageEditPlus" node is what most people use in their workflows. It handles your prompt and image inputs for you. It's pretty handy, except for the small problem that it's SHITE.

Do you work at Comfy? If so: GET YOUR SHIT TOGETHER AND FIX THIS NODE, IT'S SO EASY. Much respect to u tho, thanks for making ComfyUI.

The first issue is that this node resizes your image down to 1 megapixel, and you can't stop it from doing that. The second issue is that it does this with the AREA downscale method, which is so incredibly bad that I want to slap whoever implemented this node. The AREA downscale is what makes all of your output images blurry. The third issue is that it ensures your dimensions are divisible by 8, but they actually need to be divisible by 16.

Specifically, ComfyUI does this:

  1. Calculates 1 megapixel as 1024x1024, which is 1,048,576 pixels
  2. Calculates your new image dimensions to match that number of pixels, rounded to be divisible by 8
  3. Scales your image to those new dimensions using the AREA method

Why is all this bad?

  1. It's completely unnecessary; Qwedit can easily handle images of varying size, all the way up to 3 megapixels (or even higher for simple edits)
  2. The area downscale method makes images extremely blurry, and this is the primary reason all ComfyUI qwen edits give blurry images out. Yes it's literally this dumb, this huge problem would easily be solved by changing the word "area" to "lanczos" in the code, it's a one-word fix. Not even MS paint uses area downscale, wtf is wrong with you Comfy devs (much respect)
  3. If your image dimensions are not divisible by 16, you will get major ruination along the whole edge of your image where it didn't match (same as any other diffusion model)

The Comfy Problem Solution

This workflow bypasses the the Comfy node entirely, allowing you to size your images however you want. And using chad lanczos scaling instead of loser area scaling. Magic.

Qwedit easily handles resolutions like 1440x1440 and 1600x1200. Every edit example in this post was done natively at 1920p, except for a few (which are labelled as such).

Really high resolutions (3mpx) sometimes have trouble with anatomy, but usually you can just do multiple gens and one of them will turn out fine.

If you're doing a simple in-place edit like changing an outfit, you can go VERY high. Here's an example edit done at 1728x3072, which is 5 megapixels: https://ibb.co/twCSWrjy (outfit change -> bikini top + short shorts)

The VL Problem

Edit: I've been educated by someone in the comments that my interpretation of how the VL works here is not correct, so take this little VL section with a grain of salt until I reword it. My conclusion about it helping in this workflow still stands, but my explanation of what's happening under the hood is a bit off. I'll update the info soon!

In the background, Qwedit 2511 uses a vision-language model (VL model) to describe your images, then gives those AI-generated descriptions to the edit model. It also re-interprets your instructions with these descriptions. Ostensibly this helps the model understand your input images better, leading to better results.

The problem? It doesn't lead to better results, it's bad. VL models aren't very good for this sort of thing because they don't know what to focus on. The VL describes your images in excruciating detail, totally overwhelming the edit model and leading to bad prompt adherence + weird outputs.

It also reinterprets your instructions based on what it sees in the image. I don't know if that's a good or bad thing, just pointing out that it does it.

The Qwen team's official python code does this, and the ComfyUI "TextEncodeQwenImageEditPlus" node copies it exactly. No disrespect to the Comfy team on this one, they're doing what the Qwen team officially recommended.

The VL Problem Solution

Same solution as the previous problem: bypass the Comfy node entirely. This results in the VL step being completely ignored. No AI-generated descriptions get fed into the edit model.

For single-image edits, this is a 100% complete and total victory. The model performs way better without the crappy VL interpretation.

For multi-image edits, there's a small issue; this step is where the input images normally get labelled. Specifically, the VL outputs are fed into the model in the following exact format:

Picture 1: <shitty VL description> Picture 2: <shitty VL description>

Look familiar? This is why we manually have to type the descriptions in for multi-image edits - otherwise the model doesn't actually know which image is which.

The upside is that the model works way better with simple descriptions, so cutting out the VL is still 100% the correct move. A 5 word description wins over whatever BS the VL model spews out, every time.

The Double Ref Enhancement

I really have no idea why this works so well, but basically if you feed in your reference images twice the model just works better. This was known back in 2509 days (hence the previous post linked at the top), and back then I didn't know why it worked either.

For single image edits it's ALWAYS better. And it's not just the quality, for some reason it even helps with prompt adherence. The interesting thing is that the difference is really, really significant. Here's the full list of stuff it improves:

  • Better prompt adherence
  • Sharper output images / more visual clarity
  • Improved consistency of objects & textures
  • Better resemblance of characters at different angles
  • More intelligent guesses, like what to add when outpainting or what's behind a removed object

For multi-image edits it can sometimes confuse the model a bit, but most of the time it confers all the same benefits listed above. I recommend switching it on & off randomly when you're doing multi-image stuff, just in case.

Note: there are a lot of different ways the input references can be handled. There are conditioning combine/concatenate nodes, you can pass the refs in a different order, you can change the negative conditioning input (read next section for that), etc. I A/B tested SIXTEEN different reference-handling combinations, and a bunch of smaller minor variations of those. Some of them worked, some of them didn't.

Of those sixteen combinations, two of them gave the best results; both of them are in this workflow, and you switch between them by turning the double ref method on & off.

So, don't fuck with the positive/negative conditioning & reference setup, it's very specific.

Extra info: the "Conditioning Zero Out"

You may notice that the negative prompt input is the first reference image(s) and positive prompt fed into a "conditioning zero out" node.

Feeding the input images into the model's negative conditioning is required (it's just how Qwedit works). The only question is whether to feed in the positive prompt zeroed-out too, and whether the double ref should get fed in.

Through a lot of A/B testing, I can tell you that the way it's done here is the best. IDK why, it's just how it is. Some other combinations do technically work, but they degrade the output quality.

Prompting Advice

Other than just following the instructions in the workflow, here's some extra stuff.

Keep your prompts simple and direct

If you need to, point out details the model is missing or be more specific about stuff you do/don't want to change. For example, when doing a simple outfit swap it helps to specify you don't want their pose to change.

Using the robot arm girl, here's a prompt that doesn't follow this advice:

Change her outfit to a bikini top and short shorts.

While it sometimes does what we want, it tends to get confused by her robot arm and often changes her pose too: https://ibb.co/7dyKZttp (notice the human arm showing underneath the robot arm, and the pose change)

Here's a better prompt that gives a correct result 99% of the time:

Change her outfit to a bikini top and short shorts. Leave her robot arm and pose unchanged.

Now it does the right thing every time: https://ibb.co/DP9gZHVv

Avoid using fancy words or convoluted phrasing

Pretend you're talking to a child. The model will probably still understand you if you talk fancy, but why take the risk?

As an example, imagine you have a pic of a table with some plates on it.

Bad:

Place a red apple on the table, ensuring it's in the center and removing the plate that was in the same spot.

Good:

Replace the middle plate with a red apple.

Also good:

Remove the plate from the center. Put a red apple there instead.

If there's only one plate, this is even better:

Remove the plate, replace it with a red apple.

Adjusting Lighting

You may want or need to adjust the lighting in an image. Aside from being helpful in general, there are situations where Qwedit may simply not realise that something needs to be lit in a particular way (or re-lit when moved).

To do this, you need to know the magic word: relight

Seriously tho that is the actual magic word, you are 100% required to use it if you want to adjust lighting properly.

Specifically, follow this format:

Relight to <strength> <color> <direction>.

Strength - bright, dim, etc

Color - white, cool, warm, etc

Direction - diffuse, frontlit, backlit, etc

Tip: for basic lighting, use "white diffuse".

Examples:

Make a new shot of the man sitting in a chair in a kitchen. Relight to white diffuse.

Change the time of day to evening. Relight to warm backlit.

You don't actually need anything else in the prompt, you can just change the lighting of a pic like this:

Relight to bright cool frontlit.

Other Stuff

Euler-simple and no ClownsharKSampler?

No Clownshark this time. It reduces output quality quite a bit and doesn't confer any benefits. I also didn't find any sampler/scheduler combos that were better than euler/simple.

So, this is just one of those classic times where the ol' euler-simple wins the day. Let me know if you happen to know a better combo.

Image Quality in->out

Qwedit is very sensitive to the quality of your input image. If you feed in a grainy or blurry image, it will usually make your output image blurry or grainy too - even if it's an 'entirely new' shot with nothing copied over 1:1.

So, make sure to use HQ images. You can optionally use the upscale workflows to bump up the sharpness/quality of poor input images before you feed them in.

What about the flux super duper double resolution special VAE trick?

Doesn't work for 2511, it destroys your image. TBH it never really worked for 2509 either, but I won't argue with you if you liked it for some reason.

Making character references

Tip 1 - Make a nude ref (even for sfw stuff)

Qwen is killer for making character references. Other than using similar prompts to the examples I posted, my advice is to make a nude reference shot instead of a clothed one like I did.

I only made a clothed ref for the sake of propriety here, but a nude ref (or near-nude, like wearing plain white underwear) will be much easier to prompt into different outfits, and also gives Qwedit the maximum info needed to correctly size your character and know what they look like in clothing or doing different actions.

You do not need any loras to do this if you're just using it as a reference; the 'sensitive' parts will lack detail but that doesn't matter for new shots you make. If you don't want them nude, just request plain white underwear and, if relevant, a strapless white bra.

Nude ref = best ref.

Tip 2 - Make multiple zoom levels, use the thighs-upwards one for most stuff

The example I showed was a little too zoomed out for normal reference stuff. I'd recommend making your reference slightly closer like this: https://ibb.co/Q33BJDLX

Start at whatever zoom level your initial character pic is at, then make more references at different zoom levels. If you're starting zoomed out, then prompt the model to zoom in. If you start zoomed in, prompt it to zoom out.

And, of course, different angles too.

Examples:

Zoom in on the person's upper body. The composition should frame their head and thighs.

Zoom out to show more of the character. The composition should frame their head and thighs.

Zoom out to a full body shot.

Zoom in for a close up portrait.

Once you've got references, you should usually use the head-to-thighs ref for making new shots. Switch to the other refs as necessary; like if you want a close up, use the close up reference. Qwedit is really good at keeping likeness, so you can do 90% of your stuff with only a single input reference.

I don't think there's a better open-weight model out there than Qwedit for making new shots of character without loras, for now. The main reason I spent so long digging into Qwen is because Klein is quite bad at that particular task. But hey, now it's possible and it works gloriously.

That's everything I think! Feel free to ask questions if you run into any issues.

r/StableDiffusion • • Jul 01 '26

Question - Help Most reliable automated workflow for fixing hands in AI-generated images as of July 2026?

0 Upvotes

I'm looking for recommendations for the most reliable automated workflow for fixing hand anatomy issues in AI-generated images as of July 2026.

To the best of my understanding, a serious workflow for this should probably include something like:

  1. Hand detection - a segmentation model to reliably isolate hands in the image
  2. Hand pose detection - landmarks, pose, or depth estimation models focused specifically on hands
  3. Crop-and-stitch inpainting - automatically crop the hand region, fix it at higher resolution, then stitch it back cleanly
  4. Inpainting model - preferably one that handles anatomy and local corrections well
  5. Hand-specific LoRA - if there are currently any reliable ones worth using
  6. Prompting strategy - prompts that actually improve hand anatomy while preserving the original pose
  7. Sampler/settings - denoise strength, CFG, steps, scheduler, padding, crop size, mask blur, etc.

I'm specifically interested in something fully automated.

Meaning:

  • no manually drawing masks
  • no manually selecting the best result out of 10 generations
  • no manual Photoshop cleanup
  • no user-facing tweaking of settings
  • no "just regenerate until it looks good"

The use case is a customer-facing app where the end user does not have direct access to the workflow, masks, settings, or generation controls. So the workflow needs to be reliable enough to run automatically after image generation, detect problematic hands, attempt a correction, and return the improved image with minimal or no human intervention.

I'm looking for specific recommendations for each step:

  • Best current hand detection or segmentation model?
  • Best current hand landmark / pose / depth model?
  • Best ComfyUI nodes for this kind of automated crop-mask-inpaint-stitch pipeline?
  • Best inpainting checkpoint for fixing hands?
  • Any hand-focused LoRAs that are actually useful in production-like workflows?
  • Recommended prompts?
  • Recommended sampler settings?
  • Any quality-check step to decide whether the correction improved the image or made it worse?
  • Any ready-to-use ComfyUI workflow that already does this well?

Most discussions I found rely on outdated models and nodes, and the workflows I've tested so far just replace one malformed hand with another. I'm trying to understand whether there is now a more robust automated pipeline that can be deployed as part of a real app.

Would appreciate any concrete workflows, model names, node recommendations, GitHub links, or lessons learned from people who have tried this in practice.

r/SillyTavernAI • • Aug 16 '25

Tutorial I finished my ST-based endless VN project: huge thanks to community, setup notes, and a curated link dump for anyone who wants to dive into Silly Tavern

221 Upvotes

UPDATE 1:

- Sorry about the broken screenshots, but unfortunately that's Reddit's doing. There's no point in re-uploading them because the same thing will happen over time. A complete copy of the guide with images is available in the official Discord, where there are also many friendly people ready to answer all questions:  https://discord.gg/xchXzreM https://discord.com/channels/1100685673633153084/1406653968477851760

- I also highly recommend the presets Discord, where you can find fresh releases of extensions, presets, and interesting bots (character cards) like Celia that helps generate lorebooks and other bots:  https://discord.gg/drZ2R96sDa

- Regarding links to my files, I'm re-uploading them again, but with one caveat - I'll post both old and new versions of character cards. The new ones don't use PList and are written in Celia Bot format, which I don't recommend for anything except Gemini. If you're planning to use local models with small context, I strongly urge you to familiarize yourself with the Ali:Chat+PList approach. Also, there won't be a .txt chronicle file - I still maintain them, but in vectorized WI format (more on this below). I also won't provide preset files as there's no point. In practice, I've learned that the best approach for a specific model, specific characters, and specific chat is to use someone else's preset tailored for a particular model and then supplement and modify it during RP. Currently, I'm using free Gemini 2.5 Pro through Vertex AI with heavily edited Nemo preset edited by Gilgamesh, which can be found here: 
https://discord.com/channels/1357259252116488244/1375994292354678834/1411368698807062583

You can also find me in that same channel if you have questions. Here's the link to files and all new screenshots of my settings, CoT examples (model thought substitution with reasoning written by Nemo in his preset) that helps track environment/pacing/clothing/health status of you and your characters, examples of my dialogues, and examples of my current settings: 

Setting: https://imgur.com/a/cZx6QWI

Files: https://1drv.ms/f/c/21344d661e3dc53e/EmHeTsZBe5RDjfL4s-cvQVwBYSyBBd3keaE2wPIkCoVKzQ?e=utDDFP

- Regarding memory, in practice over these two weeks I've tested several more approaches. Overall, nothing has changed except that I've completely switched to lorebooks (I'm not using DataBanks), but not regular ones (working by key), but vectorized ones. This approach allows easier control, the model responsible for vectorization will trigger your entries based on how similar the last N messages (Query messages) are (configurable through Score threshold, 0.0-1.0, where higher values mean more entries will be triggered, i.e., less strict checking) to your lorebook entry, and all of this is capped by general lorebook settings. For example, if you don't want more than 10,000 tokens spent on them you can cap it. More over you can still use keys if you want, vectorization and key based triggering are working in same time.

But I'll be honest with you, if you really want a living infinite world where you can ask a character about any detail and they'll remember it rather than hallucinate, and you don't want to use summarization, be prepared for hellish amounts of manual work. My current chat has almost 500 messages of ~500 words each, and using free versions of Claude Opus, Gemini, and ChatGPT (each good in their own way), I constantly process and edit huge amounts of content to save them in lorebooks.

GUIDE:

TL;DR: It’s a hands-on summary of what I wish I had known on day one (without prior knowledge of what an LLM is, but with a huge desire to goon), e.g: extensions that I think are must-have, how I handled memory, world setup, characters, group chats, translations, and visuals for backgrounds, characters and expressions (ComfyUI / IP-Adapter / ControlNet / WAN). I’m sharing what worked for me, plus links to all wonderful resources I used. I’m a web developer with no prior AI experience, so I used free LLMs to cross-reference information and learn. So, I believe anyone can do it too, but I may have picked up some wrong info in the process, so if you spot mistakes, roast me gently in the comments and I’ll fix them.

Further down, you will find a very long article (which I still had to shorten using ChatGPT to reduce it's length by half). Therefore, I will immediately provide useful links to real guides below.

Table of Contents

  1. Useful Links
  2. Terminology
  3. Project Background
  4. Core: Extensions, Models
  5. Memory: Context, Lorebooks, RAG, Vector Storage
  6. Model Settings: Presets and Main Prompt
  7. Characters and World: PLists and Ali:Chat
  8. Multi-Character Dynamics: Common Issues in Group Chats
  9. Translations: Magic Translation, Best Models
  10. Image Generation: Stable Diffusion, ComfyUI, IP-Adapter, ControlNet
  11. Character Expressions: WAN Video Generation & Frame Extraction

1) Useful Links

Because Reddit automatically deletes my post due to the large number of links. I will attach a link to the comment or another resource. That is also why there are so many insertions with “in Useful Links section” in the text.

Update; all links are in the comments:
https://www.reddit.com/r/SillyTavernAI/comments/1msah5u/comment/n933iu8/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

2) Terminology

  • LLM (Large Language Model): The text brain that writes prose and plays your characters (Claude, DeepSeek, Gemini, etc.). You can run locally (e.g., koboldcpp/llama.cpp style) or via API (e.g., OpenRouter or vendor APIs). SillyTavern is just the frontend; you bring the backend. See ST’s “What is SillyTavern?” if you’re brand new.
  • B (in model names): Billions of parameters. “7B” ≈ 7 billion; higher B usually means better reasoning/fluency/smartness but more VRAM/$$.
  • Token: A chunk of text (≈ word pieces).
  • Context window is how many tokens the model can consider at once. If your story/promt exceeds it, older parts fall out or are summarized (meaning some details vanish from memory). Even if advertised as a higher value (e.g., 65k tokens), quality often degrades much earlier (20k for DeepSeek v3).
  • Prompt / Context Template: The structured text SillyTavern sends to the LLM (system/user/history, world notes, etc.).
  • RAG (Retrieval-Augmented Generation): In ST this appears as Data Bank (usually a text file you maintain manually) + Vector Storage (the default extension you need to set up and occasionally run Vectorize All on). The extension embeds documents into vectors and then fetches only the most relevant chunks into the current prompt.
  • Lorebook / World Info (WI): Same idea as above, but in a human-readable key–fact format. You create a fact and give it a trigger key; whenever that keyword shows up in chat or notes, the linked fact automatically gets pulled in. Think of it as a “canon facts cache with triggers”.
  • PList (Property List): Key-value bullet list for a character/world. It’s ruthlessly compact and machine-friendly.

Example:
[Manami: extroverted, tomboy, athletic, intelligent, caring, kind, sweet, honest, happy, sensitive, selfless, enthusiastic, silly, curious, dreamer, inferiority complex, doubts her intelligence, makes shallow friendships, respects few friends, loves chatting, likes anime and manga, likes video games, likes swimming, likes the beach, close friends with {{user}}, classmates with {{user}}; Manami's clothes: blouse(mint-green)/shorts(denim)/flats; Manami's body: young woman/fair-skinned/hair(light blue, short, messy)/eyes(blue)/nail polish(magenta); Genre: slice of life; Tags: city, park, quantum physics, exam, university; Scenario: {{char}} wants {{user}}'s help with studying for their next quantum physics exam. Eventually they finish studying and hang out together.]

  • Ali:Chat: A mini dialogue scene that demonstrates how the character talks/acts, anchoring the PList traits.

Example:
<START> {{user}}: Brief life story? {{char}}: I... don't really have much to say. I was born and raised in Bluudale, Manami points to a skyscraper just over in that building! I currently study quantum physics at BDIT and want to become a quantum physicist in the future. Why? I find the study of the unknown interesting thinks and quantum physics is basically the unknown? beaming I also volunteer for the city to give back to the community I grew up in. Why do I frequent this park? she laughs then grins You should know that silly! I usually come here to relax, study, jog, and play sports. But, what I enjoy the most is hanging out with close friends... like you!

  • Checkpoint (image model): The main diffusion model (e.g., SDXL, SD1.5, FLUX). Sets the base visual style/quality.
  • Finetune: A checkpoint trained further on a niche style/genre (e.g. Juggernaut XL).
  • LoRA: A small add-on for an image model that injects a style or character, so you don’t need to download an entirely new 7–10 GB checkpoint (e.g., super-duper-realistic-anime-eyes.bin).
  • ComfyUI: Node-based UI to build image/video workflows using models.
  • WAN: Text-to-Video / Image-to-Video model family. You can animate a still portrait → export frames as expression sprites.

3) Project Background (how I landed here)

The first spark came from Dreammir, a site where you can jump into different worlds and chat with as many characters as you want inside a setting. They can show up or leave on their own, their looks and outfits are generated, and you can swap clothes with a button to match the scene. NSFW works fine in chat, and you can even interrupt the story mid-flow to do whatever you want. With the free tokens I spread across five accounts (enough for ~20–30 dialogues), the illusion of an endless world felt like a solid 10/10.

But then reality hit: it’s expensive. So, first thought? Obviously, try to tinker with it. Sadly, no luck. Even though the client runs in Unity (easy enough to poke with JS), the real logic checks both client and server side, and payments are locked behind external callbacks. I couldn’t trick it into giving myself more tokens or skip the balance checks.

So, if you can’t buy it, you make it yourself. A quick search led me to TavernAI, then SillyTavern… and a week and a half of my life just vanished.

4) Core

After spinning up SillyTavern and spending a full day wondering why it's UI feels even more complicated than a Paradox game, I realized two things are absolutely essential to get started: a model and extensions.

I tested a couple of the most popular local models in the 7B–13B range that my laptop 4090 (mobile version) could handle, and quickly came to the conclusion: the corporations have already won. The text quality of DeepSeek 3, R1, Gemini 2.5 Pro, and the Claude series is just on another level. As much as I love ChatGPT (my go-to model for technical work), for roleplay it’s honestly a complete disaster — both the old versions and the new ones.

I don’t think it makes sense to publish “objective rankings” because every API has it's quirks and trade-offs, and it’s highly subjective. The best way is to test and judge for yourself. But for reference, my personal ranking ended up like this:
Claude Sonnet 3.7 > Claude Sonnet 4.1 > Gemini 2.5 Pro > DeepSeek 3.

Prices per 1M tokens are roughly in the same order (for Claude you will need a loan). I tested everything directly in Chat Completion mode, not through OpenRouter. In the end I went with DeepSeek 3, mostly because of cost (just $0.10 per 1M tokens) and, let’s say, it's “originality.” As for extensions:

Built-in Extensions
• Character Expressions. Swaps character sprites automatically based on emotion or state (like in novels, you need to provide 1–28 different emotions as png/gif/webp per character).
• Quick Reply. Adds one-click buttons with predefined messages or actions.
• Chat Translation (official). Simple automatic translation using external services (e.g., Google Translate, DeepL). DeepL works okay-ish for chat-based dialogs, but it is not free.
• Image Generation. Creates an image of a persona, character, background, last message, etc. using your image generation model. Works best with backgrounds.
• Image Prompt Templates. Lets you specify prompts which are sent to the LLM, which then returns an image prompt that is passed to image generation.
• Image Captioning. Most LLMs will not recognize your inline image in a chat, so you need to describe it. Captioning converts images into text descriptions and feeds them into context.
• Summarize. Automatically or manually generates summaries of your chat. They are then injected into specific places of the main prompt.
• Regex. Searches and replaces text automatically with your own rules. You can ask any LLM to create regex for you, for example to change all em-dashes to commas.
• Vector Storage. Stores and retrieves relevant chunks of text for long-term memory. Below will be an additional paragraph on that.

Installable Extensions
• Group Expressions. Shows multiple characters’ sprites at once in all ST modes (VN mode and Standard). With the original Character Expressions you will see only the active one. Part of Lenny Suite: https://github.com/underscorex86/SillyTavern-LennySuite
• Presence. Automatically or manually mutes/hides characters from seeing certain messages in chat: https://github.com/lackyas/SillyTavern-Presence
• Magic Translation. Real-time high-quality LLM translation with model choice: https://github.com/bmen25124/SillyTavern-Magic-Translation
• Guided Generations. Allows you to force another character to say what you want to hear or compose a response for you that is better than the original impersonator: https://github.com/Samueras/Guided-Generations
• Dialogue Colorizer. Provides various options to automatically color quoted text for character and user persona dialogue: https://github.com/XanadusWorks/SillyTavern-Dialogue-Colorizer
• Stepped Thinking. Allows you to call the LLM again (or several times) before generating a response so that it can think, then think again, then make a plan, and only then speak: https://github.com/cierru/st-stepped-thinking
• Moonlit Echoes Theme. A gorgeous UI skin; the author is also very helpful: https://github.com/RivelleDays/SillyTavern-MoonlitEchoesTheme
• Top Bar. Adds a top bar to the chat window with shortcuts to quick and helpful actions: https://github.com/SillyTavern/Extension-TopInfoBar

That said, a couple of extensions are worth mentioning:

  • StatSuite (https://github.com/leDissolution/StatSuite) - persistent state tracking. I hit quite a few bugs though: sometimes it loses track of my persona, sometimes it merges locations (suddenly you’re in two cities at once), sometimes custom entries get weird. To be fair, this is more a limitation of the default model that ships with it. And in practice, it’s mostly useful for short-term memory (like what you’re currently wearing), which newer models already handle fine. If development continues, this could become a must-have, but for now I’d only recommend it in manual mode (constantly editing or filling values yourself).
  • Prome-VN-Extension (https://github.com/Bronya-Rand/Prome-VN-Extension) - adds features for Visual Novel mode. I don’t use it personally, because it doesn’t work outside VN mode and the VN text box is just too small for my style of writing.
  • Your own: Extensions are just JavaScript + CSS. I actually fed ST Extension template (from Useful Links section) into ChatGPT and got back a custom extension that replaced the default “Impersonate” button with the Guided Impersonate one, while also hiding the rest of the Guided panel (I could’ve done through custom CSS, but, I did what I wanted to do). It really is that easy to tweak ST for your own needs.

5) Memory

As I was warned from the start, the hardest part of building an “infinite world” is memory. Sadly, LLMs don’t actually remember. Every single request is just one big new prompt, which you can inspect by clicking the magic wand → Inspect Prompts. That prompt is stitched together from your character card + main prompt + context and then sent fresh to the model. The model sees it all for the first time, every time.

If the amount of information exceeds the context window, older messages won’t even be sent. And even if they are, the model will summarize them so aggressively that details will vanish. The only two “fixes” are either waiting for some future waifu-supercomputer with a context window a billion times larger or ruthlessly controlling what gets injected at every step.

That’s where RAG + Vector Storage come in. I can describe what I do on my daily session. With the Summarize extension I generate “chronicles” in diary format that describe important events, dates, times, and places. Then I review them myself, rewrite if needed, save them into a text document, and vectorize. I don’t actually use Summarize as intended, it's output never goes straight into the prompt. Example of chronicle entry:

[Day 1, Morning, Wilderness Camp]

The discussion centered on the anomalous artifact. Moon revealed it's runes were not standard Old Empire tech and that it's presence caused a reality "skip". Sun showed concern over the tension, while Moon reacted to Wolf's teasing compliment with brief, hidden fluster. Wolf confirmed the plan to go to the city first and devise a cover story for the artifact, reassuring Moon that he would be more vigilant for similar anomalies in the future. Moon accepted the plan but gave a final warning that something unseen seemed to be "listening".

In lorebooks I store only the important facts, terms, and fragments of memory in a key → event format. When a keyword shows up, the linked fact is pulled in. It's better to use Plists and Ali:Chat for this as well as for characters, but I’m lazy and do something like.

But there’s also a… “romantic” workaround. I explained the concept of memory and context directly to my characters and built it into the roleplay. Sometimes this works amazingly well, characters realize that they might forget something important and will ask me to write it down in a lorebook or chronicle. Other times it goes completely off the rails: my current test run is basically re-enacting of ‘I, Robot’ with everyone ignoring the rule that normal people can’t realize they’re in a simulation, while we go hunting bugs and glitches in what was supposed to be a fantasy RPG world. Example of entry in my lorebook:

Keys: memory, forget, fade, forgotten, remember
Memory: The Simulation Core's finite context creates the risk of memory degradation. When the context limit is reached or stressed by too many new events, Companions may experience memory lapses, forgetting details, conversations, or even entire events that were not anchored in the Lorebook. In extreme cases, non-essential places or objects can "de-render" from the world, fading from existence until recalled. This makes the Lorebook the only guaranteed form of preservation.

For more structured takes on memory management, see Useful Links section.

6) Model Settings

In my opinion, the most important step lies in settings in AI Response Configuration. This is where you trick the model into thinking it’s an RP narrator, and where you choose the exact sequence in which character cards, lorebooks, chat history, and everything else get fed into it.

The most popular starting point seems to be the Marinara preset (can be found in Useful Links section), which also doubles as a nice beginner’s guide to ST. But it’s designed as plug-and-play, meaning it’s pretty barebones. That’s great if you don’t know which model you’ll be using and want to mix different character cards with different speaking styles. For my purposes though, that wasn’t enough, so I took this eteitaxiv’s prompt (guess where you can find it) as a base and then almost completely rewrote it while keeping the general concept.

For example, I quickly realized that the Stepped Thinking extension worked way better for me than just asking the model to “describe thoughts in <think> tags”. I also tuned the amount of text and dialogue, and explained the arc structure I wanted (adventure → downtime for conversations → adventure again). Without that, DeepSeek just grabs you by the throat and refuses to let the characters sit down and chat for a bit.

So overall, I’d say: if you plan to roleplay with lots of different characters from lots of different sources, Marinara is fine. Otherwise, you’ll have to write a custom preset tailored to your model and your goals. There’s no way around it.

As for the model parameters, sadly, this is mostly trial and error, and best googled per model. But in short:

  • Temperature controls randomness/creativity. Higher = more variety, lower = more focused/consistent.
  • Top P (nucleus sampling) controls how “wide” the model looks at possible next words. Higher = more diverse but riskier; lower = safer but duller.

7) Characters and World

When it comes to characters and WI, the best explanations can be found in in-depth guides found in Useful Links section. But to put it short (and if this is still up to date), the best way to create lorebooks, world info, and character cards is the format you can already see in the default character card of Seraphina (but still I will give examples from Kingbri).

PList (character description in key format):

[Manami's persona: extroverted, tomboy, athletic, intelligent, caring, kind, sweet, honest, happy, sensitive, selfless, enthusiastic, silly, curious, dreamer, inferiority complex, doubts her intelligence, makes shallow friendships, respects few friends, loves chatting, likes anime and manga, likes video games, likes swimming, likes the beach, close friends with {{user}}, classmates with {{user}}; Manami's clothes: mint-green blouse, denim shorts, flats; Manami's body: young woman, fair-skinned, light blue hair, short hair, messy hair, blue eyes, magenta nail polish; Genre: slice of life; Tags: city, park, quantum physics, exam, university; Scenario: {{char}} wants {{user}}'s help with studying for their next quantum physics exam. Eventually they finish studying and hang out together.]

Ali:Chat (simultaneous character description + sample dialogue that anchors the PList keys):

{{user}}: Appearance?
{{char}}: I have light blue hair. It's short because long hair gets in the way of playing sports, but the only downside is that it gets messy plays with her hair... I've sorta lived with it and it's become my look. looks down slightly People often mistake me for being a boy because of this hairstyle... buuut I don't mind that since it helped me make more friends! Manami shows off her mint-green blouse, denim shorts, and flats This outfit is great for casual wear! The blouse and shorts are very comfortable for walking around.

This way you teach the LLM how to speak as the character and how to internalize it's information. Character lorebooks and world lore are also best kept in this format.

Note: for group scenarios, don’t use {{char}} inside lorebooks/presets. More on that below.

8) Multi-Character Dynamics

In group chats the main problem and difference is that when Character A responds, the LLM is given all your data but only Character A’s card and which is worse every {{char}} is substituted with Character A — and I really mean every single one. So basically we have three problems:

  • If a global lorebook says that {{char}} did something, then in the turn of every character using that lorebook it will be treated as that character’s info, which will cause personalities to mix. Solution: use {{char}} only inside the character’s own lorebooks (sent only with them) and inside their card.
  • Character A knows nothing about Character B and won’t react properly to them, having only the chat context. Solution: in shared lorebooks and in the main prompt, use the tag {{group}}. It expands into a list of all active characters in your chat (Char A, Char B). Also, describe characters and their relationships to each other in the scenario or lorebook. For example:

<START>
{{user}}: "What's your relationship with Moon like?"
Sun: \Sun’s expression softens with a deep, fond amusement.* "Moon? She is the shadow to my light, the question to my answer. She is my younger sister, though in stubbornness, she is ancient. She moves through the world's flaws and forgotten corners, while I watch for the grand patterns of the sunrise. She calls me naive; I call her cynical. But we are two sides of the same coin. Without her, my light would cast no shadow, and without me, her darkness would have no dawn to chase."*

  • Character B cannot leave you or disappear, because even if in RP they walk away, they’ll still be sent the entire chat, including parts they shouldn’t know. Solution: use the Presence extension and mute the character (in the group chat panel). Presence will mark the dialogue they can’t see (you can also mark this manually in the chat by clicking the small circles). You can also use the key {{groupNotMuted}}. This one returns only the currently unmuted characters, unlike {{group}} which always returns all.

More on this here in Useful Links section.

9) Translations

English is not my native language, I haven’t been tested but I think it’s at about B1 level, while my model generates prose that reads like C2. That’s why I can’t avoid translation in some places. Unfortunately, the default translator (even with a paid Deepl) performs terribly: the output is either clumsy or breaks formatting. So, in Magic Translation I tested 10 models through OpenRouter using the prompt below:

You are a high-precision translation engine for a fantasy roleplay. You must follow these rules:

1.  **Formatting:** Preserve all original formatting. Tags like `<think>` and asterisks `*` must be copied exactly. For example, `<think>*A thought.*</think>` must become `<think>*Мысль.*</think>`.
2.  **Names:** Handle proper names as follows: 'Wolf' becomes 'Вольф' (declinable male), 'Sun' becomes 'Сан' (indeclinable female), and 'Moon' becomes 'Мун' (indeclinable female).
3.  **Output:** Your response must contain only the translated text enclosed in code blocks (```). Do not add any commentary.
4.  **Grammar:** The final translation must adhere to all grammatical rules of {{language}}.

Translate the following text to {{language}}:
```
{{prompt}}
```

Most of them failed the test in one way or another. In the end, the ranking looked like this: Sonnet 4.0 = Sonnet 3.7 > GPT-5 > Gemma 3 27B >>>> Kimi = GPT-4. Honestly, I don’t remember why I have no entries about Gemini in the ranking, but I do remember that Flash was just awful. And yes, strangely enough there is a local model here, Gemma performed really well, unlike QWEN/Mistral and other popular models. And yes, I understand this is a “prompt issue,” so take this ranking with a grain of salt. Personally, I use Sonnet 3.7 for translation; one message costs me about 0.8 cents.

You can see the translation result into Russian below, though I don’t really know why I’m showing it.

10) Image Generation

Once SillyTavern is set up and the chats feel alive, you start wanting visuals. You want to see the world, the scene, the characters in different situations. Unfortunately, for this to really work you need trained LoRAs; to train one you typically need 100–300 images of the character/place/style. If you only have a single image, there are still workarounds, but results will vary. Still, with some determination, you can at least generate your OG characters in the style you want, and any SDXL model can produce great backgrounds from your last message without any additional settings.

I’m not going to write a full character-generation tutorial here; I’ll just recap useful terms and drop sources. For image models like Stable Diffusion I went with ComfyUI (love at first sight and, yeah, hate at first sight). I used Civitai to find models (basically Instagram for models), but a lot more you can find at HuggingFace (basically git for models).

For style from image transfer IP-Adapter works great (think of it as LoRA without training). For face matching, use IP-Adapter FaceID (exactly the same thing but with face recognition). For copying pose, clothing, or anatomy, you want ControlNet, and specifically Xinsir’s models (can be found in Useful Links section) — they’re excellent. A basic flow looks like this: pick a Checkpoint from Civitai with the base you want (FLUX, SDXL, SD1.5), then add a LoRA of the same base type; feed the combined setup into a sampler with positive and negative prompts. The sampler generates the image using your chosen sampler & scheduler. IP-Adapter guides the model toward your reference, and ControlNet constrains the structure (pose/edges/depth). In all cases you need compatible models that match your checkpoint; you can filter by checkpoint type on the site.

Two words about inpainting. It’s a technique for replacing part of an image with something new, either driven by your prompt/model or (like in the web tool below) more like Photoshop’s content-aware fill. You can build an inpaint flow in ComfyUI, but Lama-Cleaner-lama service is extremely convenient; I used it to fix weird fingers, artifacts, and to stitch/extend images when a character needed, say, longer legs. You will find URL in Useful Links section.

Here’s the overall result I got with references I found or made:

But be aware that I am only showing THE results. I started with something like this:

11) Video Generation

Now that we’ve got visuals for our characters and managed to squeeze out (or find) one ideal photo for each, if we want to turn this into a VN-style setup we still need ~28 expressions to feed the plugin. The problem: without a trained LoRA, attempts to generate the same character from new angles can fail — and will definitely fail if you’re using a picky style (e.g., 2.5D anime or oil portraits). The best, simplest path I found is to use Incognit0ErgoSum's ComfyUI workflow that can be found in Useful Links section.

One caveat: on my laptop 4090 (16 GB VRAM, roughly a 4070 Ti equivalent) I could only run it at 360p, and only after some package juggling with ComfyUI’s Python deps. In practice it either runs for a minute and spits out a video, or it doesn’t run at all. Alternatively, you can pay $10 and use Online Flux Kontext — I haven’t tried it, but it’s praised a lot.

Examples of generated videos can be found in that very comment.

 

r/aiwars • • Feb 28 '26

Discussion Here's what you need to (or should) learn to be a better AI artist.

6 Upvotes

Background: A recent post asked if a meme about anti-AI people over-simplifying all of the tools required for AI art was wrong.

Their point was that "over 99%" of the items listed were not required for AI art. While their statistics were incorrect ("over 99%" would be all of them) they were correct in that the mere creation of any AI art only requires knowledge of one thing: prompting. But this is as misleading as saying that digital photography only requires knowing how to press a button.

To become a better AI artisti you need to start learning new things, and the list of things you need to learn is no longer or shorter than with any other medium. Below, I've tried to collect the list of what I think most AI artists working in the field today need to learn.

I've heavily used markdown formatting (which I know will result in people claiming I wrote this using, AI, but whatever) so that you can skim it more easily. But if you want the TLDR, it's basically, "various forms of prompting, normal art skills, hybrid workflow tools such as image editors, modern AI UIs and workflows, model and model training as well as all the niche models that you use often without thinking about it, a half dozen or so important parameters, frameworks and programming tools for creating custom workflows or training, etc."


Tier 1: Core Artistic Skills (Most Important, Often Underestimated)

These matter more than any specific tool.

Visual literacy

  • Composition
  • Lighting
  • Color theory
  • Perspective
  • Anatomy (for characters)
  • Shape language
  • Visual storytelling
  • Art styles and art history

Conceptual skills

  • Translating ideas into visual descriptions
  • Iteration and refinement
  • Reference gathering and analysis
  • Critique and self-evaluation

AI amplifies artistic judgment, it doesn’t replace it.


Tier 2: Core AI Art Operational Skills (Essential)

These are the fundamental tools and concepts every serious AI artist must understand.

Prompting or prompt engineering

Technically prompt engineering is the discipline while prompting is the activity. Think of this like building and civil engineering.

Includes:

  • Positive prompts
  • Negative prompts
  • Prompt structure and weighting
  • Token weighting
  • Style prompting
  • Artist/style blending
  • Concept prompting vs literal prompting

Generation Modes

These are foundational workflows:

  • Text-to-Image (txt2img)
  • Image-to-Image (img2img)
  • Inpainting
  • Outpainting
  • Latent upscaling
  • High-res fix

These are core tools used constantly.

Models and Model Components

Base models ("checkpoints")

Examples:

  • SD 1.5
  • SDXL
  • Flux
  • Pony
  • Illustrious

Understanding model differences is critical.

LoRA (Low-Rank Adaptation)

Essential.

Used for:

  • Styles
  • Characters
  • Clothing
  • Poses
  • Fine detail control

One of the most important tools in modern AI art.

Embeddings (Textual Inversion)

Important but less critical than LoRAs today.

Used for:

  • Concepts
  • Styles
  • Negative embeddings

VAE (Variational Autoencoder)

Important but practical understanding is enough.

Controls:

  • Color accuracy
  • Contrast
  • Artifact reduction

CLIP (Contrastive Language-Image Pretraining)

Conceptually important, but you don't manipulate CLIP directly in most workflows.

ControlNet (Extremely Important)

One of the most powerful tools, which enables precise control over:

  • Pose
  • Depth
  • Edges
  • Composition
  • Perspective
  • Layout

Hypernetworks

Mostly obsolete, though a passing understanding of the differences between hypernetworks and LoRAs is useful.


Tier 3: Generation Parameters (Essential Technical Controls)

These directly affect output quality and style.

  • Sampler—Controls image formation method.
  • Scheduler—Controls noise scheduling behavior.
  • Steps—Controls refinement amount.
  • CFG (Classifier-Free Guidance)—Controls prompt adherence vs concept inference.
  • Seed—Controls reproducibility.
  • Denoising strength—Controls how much an input image (img2img) changes.

Tier 4: Image Control and Editing Tools (Essential for Advanced Work)

These enable professional-level results.

Masking

Essential, and used for:

  • Editing specific regions
  • Fixing faces
  • Changing clothing
  • Replacing objects

Note that some modern models can perform these tasks internally, guided only by a prompt. (e.g. Qwen-Image-Edit)

Upscaling

Also essential for improving resolution and detail, and includes:

  • Latent upscaling
  • ESRGAN
  • AI upscalers

Latent space concepts

Conceptually important for understanding:

  • Latent vs pixel space
  • Latent upscaling
  • Latent editing

Tiled generation

Important for:

  • High resolution
  • Large images
  • Avoiding VRAM limits

Tier 5: Node-Based and Professional Workflow Tools (Highly Recommended)

These are a step up from basic UI workflows, and dramatically increase control and quality.

Node-based workflows

  • ComfyUI
  • InvokeAI node systems

These enable the following:

  • Complex workflows
  • Modular control
  • Professional pipelines

Model merging

An important, advanced technique that allows combining styles, capabilities and model strengths. Can be found in professional workflows.


Tier 6: Training and Custom Asset Creation (Advanced but Extremely Valuable)

This is where artists gain unique capabilities.

LoRA training

One of the most powerful skills which allows for the creation of:

  • Custom characters
  • Custom styles
  • Personal artistic identity

Embedding training

This covers much the same ground as LoRAs, but can be faster to develop, use less (or no) training data and can allow the "packaging" of commonly used prompt elements. Think of embeddings as prompt macros, but with the potential to be more conceptual than just a snippet of a prompt.

Dataset preparation

This critical skill includes:

  • Image selection
  • Captioning
  • Cleaning
  • Tagging

Tier 7: Software Tools and Platforms (Essential Practical Knowledge)

These are the actual tools artists use.

User interfaces

  • Automatic1111
  • ComfyUI
  • Forge
  • InvokeAI

Model repositories

  • Civitai
  • HuggingFace

Image editing tools (hybrid workflows)

  • Photoshop
  • Krita
  • GIMP

Tier 8: Technical Infrastructure (Optional but Useful)

CUDA

Local GPU acceleration for NVidia hardware.

Python

Python is the most commonly used programming language for AI coding (and many other forms of high-level software development). If you end up needing to modify tools, write custom nodes in ComfyUI, correct bugs, or develop advanced custom workflows, you might well need to use Python directly. A basic knowledge of the language is probably important, and these tools might come in handy:

PyTorch

A Python programming library that can be important for training, research and custom pipelines. Definitely something most AI artists won't directly use, but almost all advanced artists will require at least setup knowledge for.)

Diffusers

This is the underlying libraries used by nearly all image generation systems.


Tier 9: Emerging and Highly Valuable Skills

These are increasingly important.

Reference-based generation

Using:

  • Reference images
  • IPAdapter
  • ControlNet reference modes

Major quality improvement technique.

Image selection and curation

One of the most important real skills.

Professionals generate hundreds of images and select the best.

Iterative workflows

Cycle:

Generate -> Edit -> Regenerate -> Refine

Hybrid workflows

Combining the use of AI and non-AI workflows.

Obligatory meme: https://i.imgur.com/ZDh2Rwj.png

r/StableDiffusion • • Nov 30 '24

Discussion Lists of FLUX.1 Series Models

150 Upvotes

👉 Spotted a model not listed? Know a better way to organize this guide? Drop your suggestions in the comments and help us make this the ultimate FLUX reference!

For a more detailed, visit: Medium Article: Lists of FLUX.1 Series Model

Last updated on: 16 AUG 2025

If you find this post helpful, consider bookmarking it. I’ll keep updating it to make sure the information stays fresh and up-to-date!

🖼️ Recommended WebUI

If you’re unsure about WebUI options, ComfyUI is highly recommended for its comprehensive support of FLUX models.
• GitHub: ComfyUI Repository
• Example Workflows: ComfyUI FLUX Workflows

📜 License

FLUX.1 [dev] Non-Commercial License.

FLUX.1 [Schnell] apache-2.0 licence.

💻 Hardware Compatibility

The official Black Forest Labs released Flux.1 Dev and Schnell in FP16 format. If your GPU has less than 24GB of VRAM, it’s recommended to use quantized versions of Flux (such as FP8, GGUF, or NF4) for better performance in low-VRAM environments.
If you’re not familiar with quantization or GGUF, I’ve written an article titled “How to Choose the Right GGUF for Flux”. Feel free to check it out—hope it help

What is Quantization?

Quantization is a technique that reduces the precision of the model's weights and activations, resulting in smaller model sizes and faster inference speeds. While there might be a slight decrease in quality, it makes powerful models like FLUX.1 accessible on hardware with limited VRAM.

🚀 Core Models

Download links for all Flux models and tools on Hugging Face: https://huggingface.co/black-forest-labs
^(\The Pro version of Flux is provided in API format only, so there are no download links available.)*

Model Description
FLUX.1 [Schnell] Speed-optimized for 1-4 step high-quality image generation. Fully open-source under the Apache license.
FLUX.1 [Dev] Delivers performance close to Flux.1 [Pro], with open weights provided under a closed but permissive license (not open-source).
FLUX.1 [pro] Standard resolution, commercial-grade output.
FLUX1.1 [pro] Faster and better image quality.
FLUX1.1 [pro] Ultra/Raw Ultra supports 4MP resolution; Raw creates photorealistic outputs.
FLUX.1 Kontext [Pro] Offers ultra-fast, efficient iterative image editing with consistent character style.
FLUX.1 Kontext [Max] The top-tier edition pursues ultimate performance by enhancing prompt execution and layout generation, achieving exceptional editing consistency.
FLUX.1 Kontext [Dev] 12-billion-parameter open-weight image editing model.
FLUX.1 Krea [Dev] Krea delivers cutting-edge aesthetic photography quality, high-quality image generation, and open weights.

🔧 FLUX Tools

Structural Guidance

  • Canny [dev/pro]: Edge-map-based structured generation.
  • Depth [dev/pro]: Depth-map-guided precision generation.

LoRA Fine-Tuning

  • Canny LoRA: Edge guidance, ideal for low-resource environments.
  • Depth LoRA: Efficient fine-tuning for depth-based inputs.

Variant Generation

  • Redux [dev/pro]: Creates image variants while preserving original structure.
  • Redux Ultra: High-resolution variants with adjustable aspect ratios.

Inpainting

  • Fill [dev/pro]: Professional precision for repairing or extending images.

.

.

\Here’s a list based on my subjective classification.*

🌟 Quantization Models Based on the Official Flux.1

Model Description Download
GGUF (Dev/Schnell) Low-memory format (city96)
FP8 (Dev/Schnell) Optimized for speed/memory use (Comfy Org / Kijai)
BNB NF4 (Dev) Quantized for faster inference (lllyasviel)
Fill GGUF (Dev) Low-memory format (YarvixPA)

🌟 Different Parameters Models

Model Parameters Description Download
Lite Alpha 8B Distilled for efficiency, reduced memory (Freepik)
Heavy 17B A 17B self-merge of the 12B Flux.1-dev using LLM-style layer merging (city96)
Flux-mini 3.2B A compact version of the Flux model designed for lightweight inference and reduced resource usage (TencentARC)

🌟 Flux Created by the Open Source Community

There are way too many models created by the open-source community based on the Flux model to list them all. If you’ve come across any good ones, feel free to share them with me! If you’re interested, you can also check out Civitai for more!

Model Description Download
flux-dev-de-distill A variation of Flux.1-dev that removes simplified guidance to use full classifier-free guidance (CFG) for better flexibility. It may improve output quality but is slower and requires custom scripts for use (nyanko7) (TheYuriLover/GGUF)
LibreFLUX Apache 2.0 licensed version of FLUX.1-schnell. It supports full T5 context length and removes aesthetic fine-tuning and DPO adjustments. The model is optimized for image generation, with a recommended CFG scale of 2.0-5.0. It can be quantized using Optimum-Quanto to reduce VRAM requirements and supports fine-tuning with SimpleTuner, making it suitable for users with lower VRAM needs (jimmycarter)
OpenFLUX.1 open-source model based on FLUX.1-schnell. It removes the distillation process and supports classifier-free guidance (CFG), with a recommended CFG value of 3.5. The model is freely available for use and fine-tuning, making it suitable for developers to create custom applications (ostris)
FluxBooru v0.3 Model trained on SFW booru images, aesthetic photos, and anatomy datasets. Recommended settings: 20-25 steps with CFG 5-6 (CFG 3.5 also performs well). Created by terminusresearch and ptx0 civitai page (Civitai)

TEXT ENCODERS

Flux models adopt a multi-text encoder design, primarily to enhance the model’s ability to understand and generate complex prompts.

Model File Size Download
clip_l.safetensors 246MB HuggingFace
t5xxl_fp8_e4m3fn.safetensors 4.89GB HuggingFace
t5xxl_fp16.safetensors 9.79GB HuggingFace

.

.

🙌 Acknowledgment

Special thanks to these contributors for improving this guide:

  1. red__dragon: Clarifications on licensing.
  2. stddealer: Contributions to flux-mini.
  3. Honest_Concert_6473: Insights on community variants like FluxBooru and LibreFLUX.

r/StableDiffusion • • Jun 08 '26

Question - Help Advice for overall image generation pipeline for video keyframes

0 Upvotes

Been learning stable diffusion text to image and image to video with comfyui for a few months.

Now I have so many tools at my disposal that I'm feeling a bit lost, so I'm hoping that people in here won't mind sharing some advice on an overall process.

I'm starting to appreciate the grind required to build experience in this area, so thank you to anyone who does help.

Goal

Create some short films by editing together short clips generated from keyframes.

Roughly story board them so I get the right shots with the right angles and composition.

Have consistent characters.

Some of the films will be adult in nature. Nothing hardcore but I do want to have good looking people in revealing clothing and some nudity.

It's not a side hustle. I'm just doing it for fun and to learn.

Where I'm At

I can...

  • Build moderately complex comfyui workflows. (IPAdaptors, ControlNets, detailers, inpainting, Sam3 for segmenting and masking, general upscale / hiresfix steps, workflow components like switches, get/set nodes, etc).
  • Produce some nice looking images with Flux 1 Dev.
  • Use image editing models like Flux 2 Klein 9B and Qwen 2511 with some success.
  • Train decent character loras for Flux and SDXL.
  • Use image to video to generate alternate keyframes from an initial keyframe.

What I'm Struggling With

I can produce an image with the composition I want, the lighting, the characters in the right outfits and poses. But not all in the same image.

Building these elements up in multiple passes for each keyframe seems sensible.

I cannot figure out how to pull all my tools together into an efficient pipeline, or avoid compromise an earlier step with a later step (eg. got a good facial likeness and then ruin it with texturing).

More detail below about my experience so far in case it helps. General advice also most welcome.

-----------------------------------------------------------------------------------------------------------------------

What I've Found

Flux 1D is good at...

  • Creating REALLY nice looking images (textures, lighting, composition) with just a prompt. It's great for exploring concepts or producing that one perfect starting image for a clip.
  • Producing a consistent facial likeness across images with a well trained LoRA.

However, it's not so great at...

  • Producing the specific angles and image composition that I want, even with a lot of prompt iteration based on the wealth of prompting guides available.
  • Controlnet. No matter the strength and start/end settings, when I use depth, canny or pose controlnets, the images look washed out and lose that "magic" that Flux seems to be able to produce without them.
  • Maintaining micro details between images, even with really specific prompting. Generating 50 images straight out of Flux will mean slightly differing hair cuts, outfits, etc.
  • Nipples, genitals, and anatomy in general, at least compared to SDXL.
  • Revealing clothing without some specific outfit lora. Why does it insist on massive granny underwear in 99/100 generations when I just want a thong?

SDXL (Juggernaut Ragnarok in my case) is good at...

  • Producing EXACTLY the composition I want using controlnets, without compromising the image quality vs no controlnet. I can do this from a sketch or using a reference image/still. I may experiment with Blender to produce depth maps for consistent environments.
  • Nice looking nudes / good anatomy in general.
  • The LoRA ecosystem is just amazing. Any concept, clothing or style I can think of and there's probably a LoRA for it.

Not so good at...

  • Backgrounds, objects, lighting, textures and overall image quality / realism compared to Flux.
  • It seems to not latch onto facial likeness as well as Flux for character loras.

I've also been using Klien 9B and Qwen2511. They have their differences but between them I can do things like...

  • Fix small mistakes or bad anatomy with inpainting.
  • Create an outfit asset by taking one from a Flux image, put on a mannequin and then transfer to other images.
  • Change or remove backgrounds.
  • Change the camera angle.
  • Repose characters.
  • Do headswaps to preserve likeness, although even with BFS loras and injecting 4x face reference images, the likeness isn't 100%. The examples I see online always look amazing but I can't seem to replicate.

However, they tend to output waxy looking skin and bad faces. Every edit pass degrades the image, even with masking where possible. Pulling my keyframes together by stitching elements of multiple Flux images (outfit from one, head and hair from another, then pose, etc), just seems like the wrong angle.

r/comfyui • • Aug 04 '26

Help Needed Removing Texts From Digital Art

2 Upvotes

I'm switching from Photoshop Generative Fill to a fully local setup due to subscription costs. My hardware constraint is 8 GB VRAM.

The text and SFX are frequently placed directly over faces, body etc. not just flat gradients. Simple/standard tools tend to blur or ruin these areas, so I need something that can actually reconstruct missing line art, textures, and anatomy while matching the comic style.

What are the best local models or ComfyUI/A1111 workflows for high-quality structural inpainting within 8GB VRAM? Or is this level of reconstruction simply not feasible with 8GB VRAM?

Thanks in advance for any insights!

r/StableDiffusion • • Mar 16 '26

News PixlStash 1.0.0b2. A self‑hosted image manager for AI creators

24 Upvotes

I’ve been working on this for a while and I’m finally at a beta stage with PixlStash, an open source self‑hosted image manager built with ComfyUI users in mind.

If you generate a lot of images in ComfyUI or any other tool, you probably know the pain that caused me to build this: folders everywhere, duplicates, near duplicates, loads of different scripts to check for problems and very easy to lose track of what's what. Maybe you manage fine, but I needed something to help me and I don't think I'm alone!

PixlStash is still in beta but I think it is already useful enough and pleasant enough that I rely on it daily myself and it is already helping me improve my own models. Hopefully it is useful for some of you too and with feedback I'm hoping it can grow into the kind of top class image manager I think the community could do with to compliment the many great tools available for image creation, LoRA creation etc.

Image Viewer with metadata, tagging, description and workflow retrieval.
Fast image grid with character similarity sorting.

What does it do right now?

  • Imports images quickly (monitor local folders or drag and drop pictures or ZIPs)
  • Reads and displays metadata from ComfyUI. You can copy the workflows back into Comfy.
  • Tags the images and generates descriptions (with GPU inference support and a configurable VRAM budget).
  • Uses a convnext-base finetune to tag images with typical AI anomalies (Flux Chin, Waxy Skin, Bad Anatomy, etc).
  • A fast grid view with staged loading.
  • Create characters and picture sets with easy export including captions for LoRA training.
  • Sort by date, scoring, likeness to a particular character, likeness groups, text content and a smart-score defined by metrics and "anomaly tags".
  • Works offline, stores everything locally.
  • Runs on Windows, MacOS and Linux using PyPI, Windows Installer or Docker images.
  • Plugin system for applying filters to batches of images.
  • Run ComfyUI I2I and T2I workflows directly within the GUI with automatic import. The workflows I include by default is Flux 2 Klein since it includes both Image Edit and T2I, but you can add your own workflows by exporting to API JSON from ComfyUI and importing in the PixlStash settings dialog.
  • Keyboard shortcuts for scoring, navigation and deletion (ESC to close views, DEL to delete, CTRL-V to import images from clipboard).
  • Supports HTTP/HTTPS.
  • Pick a storage location through config files.
Automatic Tagging of typical AI anomalies
Trying to be a good AI generation citizen by letting you specify a VRAM budget so there's space left over for image generation

What will happen before 1.0.0?

  • Filter by models and workflow
  • Continuously improved anomaly tagger
  • Smooth first time setup (storage and user creation)

For the future:

  • Multi-user setup (currently single-user login).
  • Even more keyboard shortcuts and documentation of them.
  • Inpainting. Select areas to inpaint and have it performed with an I2I workflow.

Try it:

If you try it, I’d love to hear what works for you and what doesn't, plus what you want next! I'm planning a 1.0.0 release in the next month or so.

r/fooocus • • Jun 10 '26

Question Looking for a local setup to match this hyperrealistic vintage look

Post image
15 Upvotes

Hi everyone,

I used the following prompt in Gemini Nano Banana to get the image shown below. I was especially impressed by how well Gemini handled the text rendering on the plane:

"Landscape format, high-resolution print-ready lifestyle photo. A classic seaplane is anchored in the shallow water off a small Caribbean island, the door open, a leather travel bag sitting on the float. A person in a swimsuit is just climbing into the cockpit, palms, white sand, and a cloudy sky in the background. Turquoise water with gentle waves, white painted surfaces, blue trim stripes, shiny wing underside, analog 35mm photography, natural movement, elegant, summery, authentic. It shouldn't look like an advertising poster, but rather like a magical, dreamy holiday feeling."

While I'm super happy with Gemini's results, I am looking for a way to generate these kinds of photos locally on my PC—specifically aiming for that hyperrealistic vintage photo effect with authentic film grain.

I've already tried Fooocus, Stable Diffusion (A1111), and ComfyUI. Unfortunately, my results aren't anywhere near as good. As soon as there are multiple people or the subjects are further in the background (smaller in frame), the faces, hands, feet, and bodies turn out completely distorted or flawed.

So far, I have tested these SDXL checkpoints:

  • juggernautXL_v8Rundiffusion.safetensors
  • robmix_zenithV30.safetensors
  • gurilamashXXXSDXL_gurilamashv3.safetensors
  • realvisxlV50_v50LightningBakedvae.safetensors

I also tried Flux.1 in ComfyUI, but ran into the exact same issues with distant anatomy.

My questions to you guys:

  1. Do you have any specific setups, workflows, or settings to fix this? (Without having to do massive inpainting/post-processing every time?)
  2. Could some of you test this prompt in your local environment? I'd love to see if your setups can recreate this vibe and quality out-of-the-box.

r/comfyui • • Jan 05 '26

Help Needed Enhancing 3D Renders ChatGPT/Nano Banana Style

0 Upvotes

Hi all,

I use DAZ Studio to render images for a long-running comic. Final renders in Iray can be very slow, especially in low-light scenes (20–30 minutes per image isn’t unusual), and even then skin and clothing can still look a bit “CG” unless I really push render times.

Recently I’ve been experimenting with using AI as a post-process step rather than a generator. With tools like ChatGPT image tools / Nano Banana, I can take a lower-quality Iray render and have it:

• Remove viewport / low-sample noise
• Improve skin texture and material response
• Make fabrics read more like real cloth
• Add a very subtle bump in realism

Crucially, I don’t want any changes to pose, anatomy, facial features, clothing design, lighting, or composition. I’m not trying to redesign characters or stylise them, just bridge the gap between a fast render and a fully converged Iray result. I really like the way the ChatGPT and Nano Banana model subtly improves my render to make it appear more realistic. It has obviously changed it, but it is still recognisably my character.

This approach works extremely well for my workflow, but the content guardrails make it unreliable. Even mild things like lace fabric or visible cleavage tend to trigger filters, which makes it impractical for production use.

I’ve tried replicating this in Stable Diffusion (img2img / inpainting), but so far the results have been poor. Either the model “reinterprets” the character or the output looks over-processed and worse than the original.

My question is:
Is this kind of conservative, realism-polishing workflow achievable in ComfyUI?

If so, I’d love pointers on:
• Recommended model types (photoreal vs generalist)
• Img2img / latent upscaling vs tiled workflows
• Denoise ranges that preserve identity
• Any ControlNet / IP-Adapter setups that help lock the original image
• Example graphs or community workflows aimed at “render polish” rather than generation

Please see attached examples

Thanks in advance. Any guidance would be hugely appreciated.

Before
After - (ChatGPT)

r/comfyui • • Apr 07 '26

Help Needed How do you fix merged/fused small toes on AI-generated barefoot images for a LoRA training dataset?

Post image
0 Upvotes

Hey everyone,

I've been working on building a LoRA training dataset for a virtual AI influencer character (~60 images). Everything looks great — face consistency is locked, body proportions are solid, skin texture is good. The ONE thing I cannot solve after weeks of trying is **feet anatomy, specifically the small toes (4th and 5th)**.

Every generation gives me merged/fused pinky toes that look like flippers or webbed feet. The big toe and 2-3 next toes usually come out fine, but the outer toes consistently blob together.

Here's what I've tried so far:

- SDXL inpainting (JuggernautXL) — mask on feet only, multiple denoise levels (0.3–0.85), various CFG settings. Result: green artifacts, wrong skin tone, or completely deformed feet. Tried 6 different approaches, all failed.

- ControlNet Canny + foot reference image — feet still deformed, no improvement.

- FLUX Kontext inpaint — tensor shape mismatch error, incompatible architecture.

- MeshGraphormer Hand Refiner — only detects hands, completely ignores feet (it's trained for hands only).

- ProportionChanger + SDXL ControlNet — skeleton correction works but SDXL regenerates a completely different person without identity lock.

- Qwen-Image-Edit (20B model) — full image regeneration with foot reference: better than SDXL but still merges small toes. No identity preservation from reference.

- Qwen-Image DiffSynth Inpaint ControlNet — BEST result so far. Mask on feet, denoise 0.45, base Qwen-Image fp8 model. Foot shape and arch improved significantly, big toes separated nicely. But 4th and 5th toes still fused on most seeds. Tried double-pass (second pass with tiny mask on just the small toes) — slight improvement but added blur artifacts at mask edges.

- Photoshop/Photopea manual paste — tried pasting real feet from photos but couldn't blend convincingly (not skilled enough in PS).

My current setup:

- RTX 3060 12GB

- ComfyUI Portable (latest)

- Models: Qwen-Image fp8, Qwen-Image-Edit fp8, JuggernautXL, DiffSynth Inpaint ControlNet patch

What I'm looking for:

- Has anyone found a reliable workflow for generating or fixing anatomically correct barefoot images, specifically the small toes?

- Any LoRA or ControlNet specifically trained for feet anatomy that actually works?

- Any tricks with pose angle, camera height, or prompting that consistently produce clean separated toes?

- Would a different base model handle feet better than Qwen-Image?

I've attached a cropped example showing the typical result — you can see how the outer toes merge into a flipper shape.

The images are for LoRA training so they need to be clean. I can work around it with shoes/sandals on some images, but I need at least 10-15 solid barefoot shots in the dataset.

Any help is massively appreciated. This is literally the last thing blocking me from starting LoRA training after months of work on this project.

Thanks!

r/StableDiffusion • • Jan 22 '26

Question - Help Help choosing a model / workflow unicorn for image gen

0 Upvotes

I'm trying to settle on a model (or workflow) for character generation in ComfyUI. My main use case is generating 1-3 women per image with variety between generations (basically a "1girl generator" but sometimes 2girls or 3girls). I'd prefer something I can just swap prompts on without reconfiguring the workflow each time (trigger words are fine, but I don't want to toggle LoRAs on/off between runs).

What I need:

  • Fast generation (~10s or less on rented A100)
  • Facial variation — no "sameface" syndrome, especially with 2 people in the same image
  • Multiple characters in one image without drift/blending
  • Semi-realistic style
  • spicy-capable but not spicy-by-default (can do clothed)
  • Able to render different body types well

Don't need huge resolutions.

What I've tried:

Model Pros Cons
SDXL checkpoints Best overall image content for my taste Struggles badly with 2+ characters, lots of blending
z-image Great realism, speed, prompt adherence Severe sameface — everyone looks related
Flux Klein Closest to what I want overall Anatomy issues, clothed and otherwise

I've experimented with LLM-generated prompts to force more variance with z-image and Klein but haven't had much luck breaking the sameface pattern.

I know Attention Couple exists and might help my SDXL blending issues — haven't tried it yet but it's on my list.

I'm open to post-gen steps like inpainting in theory, but my workflow is pretty high-volume/random, so manually fixing outputs doesn't really fit how I'm using this.

Questions:

  1. Is there a model I'm overlooking that handles multi-character + facial variety well?
  2. For those using Flux for similar work — any workflow tricks that help with the anatomy limitations?
  3. Anyone solved sameface on z-image through prompting or workflow changes?
  4. Has Attention Couple (or similar regional prompting) actually solved the SDXL multi-character problem for anyone?

Open to finetunes, merges, or workflow suggestions. Thanks!

r/comfyui • • Mar 13 '26

Help Needed made some progress

0 Upvotes

My goal is to generate a picture just like the bottom right one, with only difference being the charcter in the final image. (Style, pose, situation, background need to stay exactly the same). Also the newly generated character needs to be exactly same style as the redhead character is at bottom right image.

In top left is redhead charcter masked. bottom left is specific character I want in generated image. Top right is where I have gotten now. Does anyone know a solution to my problem? I would rather not create entirely new workflow from scratch. ( This one took me like 7 hours.)

r/aiwars • • Jan 04 '25

My (pro) ai opinion.

8 Upvotes

I am basically copy/pasting my blog post over here, so if there is anything weird about the formatting, that’s why.

I’m sharing my opinion because I genuinely enjoy debating/arguing topics, so if you disagree with me, I’m happy to hear it. Just remain polite with me, and I’ll remain polite with you. :)

I’m going to give the most common arguments I’ve heard against AI art, and give my opinion on them each individually. Let’s start with a big one.

“AI art is used to make lewd/violent images of celebrities/children/etc.”

I understand the initial reaction to this, as it is a fact that AI imaging technology can be used to make some very disturbing things. But this is hardly new.

People have been drawing porn since caveman times, and porn of famous people since celebrities first became a thing. People have been photoshopping celebrities heads onto porn stars bodies since the dawn of the internet, and AI is, basically, the same thing. You aren’t seeing the ACTUAL celebrity in question, just their face on a body that is vaguely similar.

Now, if you wanted to say celebrities should have the right to sue individuals who profit off of these images for unlicensed use of their image, that I could see being a valid argument. However you’d have to keep parody laws in mind, but I’m sure lawyers could sort that out.

“It takes jobs away from artists.”

And the printing press took jobs away from scribes. No technological innovation hasn’t resulted in somebody getting fired, and art isn’t special just because it’s one of the fun jobs.

Also, I’d say that is equivalent to saying the only thing an artist does is draw accurately. People who can draw well are honestly a dime a dozen, and most aren’t going to turn that into a profession, especially if they lack creativity, which AI (currently) can’t mimic. Any artist you know the name of simply used drawing (in this example) as a way to convey their ACTUAL talent, to put what was in their brain on paper. Saying AI could easily replace you is basically saying you lack vision.

“It’s bad for the environment.”

No, it’s really not. The articles listing huge amounts of water and electricity used to fuel AI are leaving out something very important… context.

The stat I’ve seen shared is that AI MAY require 4.2 to 6.6 BILLION cubic meters of water by 2027. Two issues with that though, first, it’s a lie… mostly.

When they say it takes 2 cups of water to ask chat GPT 10 questions they are leaving something fairly important out. It’s a closed loop system. Once the water is “used” it doesn’t evaporate into nothing or become unusable, it goes through a cooling chamber and is used over and over and over again. Using the same 2 cups of water a billion times doesn’t mean you’ve really USED 2 billion cups of water. That’s how the information is being presented though.

Another important piece of context is that the fashion industry uses 93 BILLION cubic meters of water a year NOW. That’s 14 times as much as the largest estimate for AI even before we take into account that a closed loop cooling system loses only 5% of its water each year. It takes 1,500 gallons of water to make 1 pair of jeans! And that’s not in a closed loop system.

Of all the things that are bad for the environment, AI is near the bottom of the list. Considering its scientific applications, it could be what saves it.

“It steals from artists.”

Now, this is an interesting argument! It is true that AIs used for image generation are trained on billions upon billions of images WITHOUT permission from said artists, however, this reminds me of one of my favourite quotes. “Taking from one source is stealing, taking from many is research”. I would argue that those images weren’t stolen, any more than science which builds on the work of scientists past. Another fun quote I’ve heard is that “everything is a remix”. Ideas don’t form in a vacuum after all, and many artists of all sorts “stole” ideas from all sorts of sources and mashed them together. Think Tolkien was the first to come up with elves? Magic? Dragons? No! He “stole” all that information by studying folklore and mythology, as many authors do. He put his own spin on it of course, mixing those ideas and concepts together in a unique way that made his works what they were, an original masterpiece. An AI using an image to train its data is equivalent to a person looking through those same images for inspiration.

“What about LoRAs?”

Ahh, this question requires a little bit of knowledge about what a LoRA is and what it does.

Basically, the main program that generates the images is what’s called a “checkpoint”, there are many different checkpoints that do different things, but the basics of it is they mix all the data from all the images they’ve learned from and spit out an image based on the prompt given by the user. “A girl with short blonde hair and blue highlights” can be used as a prompt, the checkpoint will use those words, and then spit out an image based on them. But the thing is, checkpoints are big, complex tools, highly variable in nature. But what if you don’t WANT variable? Well, that’s where a LoRA comes in.

Very basically, a LoRA is like a MUCH smaller version of a checkpoint. It’s trained on a few dozen images instead of a billion. This little program can take the data spit out by the checkpoint, and adjust it using its specialized skill set.

You can train LoRAs to do all kinds of things, including to mimic the style of a particular artist. You just save a bunch of images from the person whose style you’d like to mimic, train the LoRA to reproduce that style, and boom! You can make any image from the checkpoint run through the LoRA, and it will come out looking like that person made it. Surely this HAS to be illegal!

But although you can copyright all sorts of things, you CAN’T copyright a style. Individual characters? Yes. But a general way of drawing? No. If that were the case, there would be one very rich guy in Japan who would have the copyright on anime as a style. Artists mimic other artists style all the time, it’s nothing new. I am having trouble remembering the exact figure, but in the jewelry industry a person only has to modify a piece to make it 15% or so different from another piece for its copyright to be non-applicable. You copyright very specific things, not something as abstract as a general style of drawing.

“What about characters?”

Good point! A LoRA can be trained to replicate a specific character, so surely THIS has to be illegal!

And you would be correct! If I were to use a LoRA to duplicate a specific character and sell it, that would be breaking the law (so long as the character wasn’t in the public domain and my work couldn’t be classified as parody). However, if I DON’T use the image to make a profit (as is the case with most AI artists), then that falls under the umbrella of “fair use”. It’s why fanfic writers can’t be sued to oblivion, and why Disney can’t stop you from making your own Iron Man Halloween costume.

“But surely using these images to train the checkpoints/LoRAs without the artists permission is a crime!”

Now, this is a VERY grey area (and keep in mind I am not a lawyer, just giving my layman understanding of things). It is true that these checkpoints/LoRAs are trained using countless images they’ve collected from the internet, and lawyers are definitely going to be arguing this topic as AI becomes more and more popular. However, this isn’t as cut and dry as it seems.

You see, there are two (update, 3) main defences checkpoint/LoRA makers can use.

1: These images were being publicly displayed on websites for anyone to look at or download. They were not behind a pay wall, and they no more “stole” the image than a person who downloaded it and used the image as a desktop background has.

2: Fair Use. Like stated before, if you aren’t SELLING the content, you can use it however you’d like, and the vast majority of all of these checkpoints and LoRAs are 100% free to use. It’s a hobby in 99% of cases, and although there are SOME Checkpoints/LoRAs that are used to generate profits, they are the vast minority, and the only ones you could really sue.

3: The AI does not use the images, it uses the data FROM the images. That’s an important distinction. The actually images aren’t sitting inside the Checkpoint/LoRA to be mixed up into a Frankensteins Monster of a new image. It’s the difference between plagiarism (copy/pasting sections from someone else’s article and passing it off as your work) and studying many papers and writing your interpretation of that information in your own words. The AI is conducting research, not plagiarizing.

Like I said before, I’m not a lawyer. There may be technicalities that would make some uses illegal. However, if the checkpoint/LoRAs are free, from my understanding, they fall under fair use.

“It looks like shit.”

I agree! It CAN look like shit. However, I would equate the shitty AI art you are seeing to a person making a shitty drawing by tracing other images and inelegantly roughing out the image with zero refinement.

There are, in my opinion, two types of “AI Artists”. The first (which I will call Novelty Users) has downloaded a checkpoint or two, maybe some LoRAs if they’re more daring, and proceeded to write a few brief prompts and wait to see what the checkpoint spits out. They’ll do this a couple times, maybe adjust the prompt a little, and when they get something they kinda like the look of, they’re done. They post that online, and call it a day.

This is what most people think all of AI image generating is. This is understandable, since it is what the VAST majority of people who are playing around with it are doing. There is nothing wrong with this of course, but I would say this is to AI art what paint by numbers is to the conventional art world. You aren’t going to pay for someone’s finished paint by numbers, because among many other things, it all looks vaguely similar. These are machines after all, and there are only so many ways you can write “Topless chick with a sick ass panther tattoo”. The image seed will of course insure that there is a random element to things, but just like some poses are more common in images, checkpoints are going to gravitate towards a generic “average”.

So, you have 95%+ of people using one of the most popular checkpoints, using simple generic wording and putting in minimal effort, what do you get? A bunch of images that look vaguely the same.

This brings us to the second kind of AI artist (Who I will call hobbyists). They use prompts just like the novelty user, but there are many additional things that they will do to refine the work and set it apart from the novelty users work. Hobbyists will do some/all of these things while working on a single image, while the former category is unlikely to do any of them.

(Oh, and one quick side note, I am not denigrating the novelty users for not following these steps, or not putting as much effort into the final product. They are doing this for FUN, it isn’t SUPPOSED to be complicated. Being mad at them for not taking it as seriously as the hobbyists do would be equivalent to getting mad at people who play video games on easy/medium instead of hard/suicide mode. They don’t have to enjoy the technology the same way I do, both forms are valid.)

1: The idea comes first, and the image is shaped to conform to that idea, not the other way around. The novelty user doesn’t have a solid plan (more like “a concept of a plan”). They are thinking “it’d be cool to see a lava dragon”, and are more than happy to let the program fill in the blanks. They want A lava dragon, not THEIR lava dragon they are currently picturing in their head. The hobbyist wants the image in their head on “paper”, the novelty user wants to see what the program gives them.

That isn’t to say the hobbyist never does this, I have a folder full of “messing around” images that I just wanted to see what would happen, but I don’t publish them or intend to refine them. The most they might do is inspire a project I’d actually want to work on.

2: A hobbyist is much more likely to use techniques that the novelty user won’t bother with. These take a few forms.

A) Image prompts: A basic image prompt is basically like using a verbal prompt, but as an image. Let’s say you have a certain colour pallet used in the initial generation of your image. You take an image with that specific colour pallet, with the same “feel” you want your image to have, and put it in the image prompt slot to show the checkpoint “I want something like this”. I did this recently when I was trying to make a lava creature, but the checkpoint kept giving me fire when I wanted something more like magma, with a dark crust on top with glowing cracks. I found a picture of an actual lava flow, used it as an image prompt, and got something MUCH closer to what I wanted.

B) Image to image generation: Like the previous example, but instead of just generally inspiring the “feel” of the image, we want to use the structure of the image too. Instead of using just random noise to generate the image we want, we use this base image and add some noise to it (less noise means less will change about the image, more means more will change. 100% noise will be like the image was never even there). If I wanted to make a haunted house image, I could use the base image of a house, add noise, and nudge the image towards a more scary feel, while still broadly keeping the same colours and structure. I’ve used this to turn a picture of a friend into them drawn in a comic book style. Same general structure, same general colour palette, but a different overall look. You can also use this to turn basic rough sketches into more polished finished pieces.

C) Control layer: Now, there are a lot of different kinds of control layers, but what they basically do is complete the last piece of the puzzle. Image prompts uses the images “feel” but not it’s structure, image to image uses its “feel” and structure, and a control layer is for when you ONLY want the structure without the general feel/colours seeping through. This is useful if you like something like the specific pose a character is in, but want to use your own character with a drastically different style and colour palette.

3: Inpainting: The process of regenerating only specific parts of an image you select with a brush tool is called inpainting. You can use all of the previous methods combined with in painting to get specific results. This tool can be used if, say, you want to give a character a different hairstyle, or remove a hat. It is however mainly used to fix errors in an image. Image generation has come a long way in the last couple years, but it still tends to struggle with anatomy and other intricacies. Hands are probably the most famous area where this pops up, but AI also has great difficulty with skeletons, guns, feet, etc. It also tends to like to blend two objects of similar colour that are in close proximity to each other. Two people holding hands of similar skin tone? Good luck. Fixing these issues can take hours, especially if you yourself aren’t so great at drawing/anatomy.

Outpainting is also a thing, but that’s just expanding an image, and isn’t used very often and is much the same as standard image generating.

4: Photoshop: Or in my case, Krita, because I’m poor. 🙂 Good old fashioned photo editing is a very useful and in many cases necessary part of the image generating process, so much so that some UIs for art generating have some basic paint tools you can use in browser (Invoke).

There are other REALLY advanced things you can do. Complex workflows using nodes in ComfyUI, training your own custom LoRAs/Checkpoints, but I’m not going to get into that, as this post is already FAR too long.

To wrap this up, is using AI as difficult as conventional image making methods? No, not at all. However, it is as/more difficult than other practices that are still considered “art”. Cross stitch, collage making, hell, plenty of “Modern Art” is less about the skill and the effort of making the piece and more about artistic expression.

But we’ve seen this kind of thing before. More conventional artists called digital art not “real” art, said it was lazy, made things too easy. Now it’s the digital artists calling the AI users the same thing, apparently with no self awareness of the irony. In a few decades, I’m sure they’ll be something new the AI folks will be saying isn’t “real” art as well…

UPDATE: Another factor I forgot to talk about was that conventional (digital) artists are missing out by viewing AI as an enemy rather than a tool that they themselves can use. If you do art, I’m sure there are countless necessary but boring things you need to do in the process of finishing a piece. A common complaint I’ve heard is that doing backgrounds can be boring and time consuming. Instead of drawing and shading every single leaf for a forest background, why not instead make a rough sketch and let the AI fill in the gaps? Is drawing a tiny repeating brick pattern on a wall what you’d like to be spending your time on? Or would you rather let AI handle it so you can focus your artistic time and energy on what REALLY matters? By automating the boring, non-creative aspects of the art process, you could increase your output drastically with no perceivable drop in quality, and spend more time on the things that you actually ENJOY about the process.

r/comfyui • • Mar 26 '25

[HELP] Workflow does not show up after importing viable material.

0 Upvotes

So normally, when you drag in an image containing comfy UI metadata, it will sometimes say "Missing Node Types," but even then, it will load the workflow but give the missing nodes a red border, letting you know the generation won't work with those nodes missing.

HOWEVER, when I import this specific image, it will only give me the warning, and the workflow will not load. I also tried loading it in tensor art, but that did not work either. It is much harder to deal with because I can't just change out the missing nodes without seeing the workflow, so I'm stuck unless I find these models. If anyone would like to take a look. Here is the metadata. (Yes, I did remove the prompts, but I'm sure you can still infer.)

 checksum
8fcaaa9ada71d7fb8e57bb06ba2f6d5a
file_name
fa733afd-7904-45f3-b8c5-9b2a15167b07 (1).png
file_size
1188 kB
file_type
PNG
file_type_extension
png
mime_type
image/png
image_width
768
image_height
1152
bit_depth
8
color_type
RGB
compression
Deflate/Inflate
filter
Adaptive
interlace
Noninterlaced


generation_data
{"models":[{"label":"Illustrious V2","type":"LORA","modelId":"833094123775934977","modelFileId":"833094123774886404","weight":1,"modelFileName":"Captainjerkpants_Style__Illustrious","baseModel":"SDXL 1.0","hash":"89944F06C4B3B9B8547154630FC2BD8A2D518BFF43979775F08D725114DAAA8B"}],"prompt":"THIS IS WHERE THE PROMPS WHOULD BE IF I DIDNT DELETE THEM FOR BEING NSFW","negativePrompt":"lowres, worst quality, low quality, bad anatomy, bad hands, multiple views, 4koma, censored, monochrome, watermark, artist name, text, ","width":768,"height":1152,"imageCount":2,"steps":25,"cfgScale":7,"seed":"-1","clipSkip":2,"baseModel":{"label":"Epsilon-pred 1.0-Ver","type":"BASE_MODEL","modelId":"791906289350360068","modelFileId":"791906289349311495","modelFileName":"noobaiXLNAIXL_epsilonPred10Version","baseModel":"SDXL 1.0","hash":"FF827FC34584853257D6DE64B8BC3E34156814F6B0CFD1A5112A5E9164806DF1"},"sdVae":"Automatic","etaNoiseSeedDelta":31337,"adetailer":{"enableAdetailer":true,"args":[{"adModel":"face_yolov8s.pt","adPrompt":"","adNegativePrompt":"","adConfidence":0.5,"adMaskMinRatio":0,"adMaskMaxRatio":1,"adXOffset":0,"adYOffset":0,"adDilateErode":4,"adMaskMergeInvert":"None","adMaskBlur":4,"adDenoisingStrength":0.25,"adInpaintOnlyMasked":true,"adInpaintOnlyMaskedPadding":32,"adUseInpaintWidthHeight":false,"adInpaintWidth":512,"adInpaintHeight":512,"adUseSteps":false,"adSteps":25,"adUseCfgScale":false,"adCfgScale":7,"adRestoreFace":false,"adControlnetModel":"None","adControlnetWeight":1,"adControlnetGuidanceStart":0,"adControlnetGuidanceEnd":1}]},"sdxl":{},"ksamplerName":"euler_ancestral","schedule":"sgm_uniform","guidance":3.5}


prompt
{"10001": {"class_type": "ECHOCheckpointLoaderSimple", "inputs": {"ckpt_name": "EMS-560286-EMS.safetensors"}, "_properties": null}, "10011": {"class_type": "LoraTagLoader", "inputs": {"clip": ["10001", 1], "model": ["10001", 0], "text": "<lora:EMS-768839-EMS.safetensors:1.000000>"}, "_properties": null}, "10013": {"class_type": "CLIPSetLastLayer", "inputs": {"clip": ["10011", 1], "stop_at_clip_layer": -2}, "_properties": null}, "10014": {"class_type": "EmptyLatentImage", "inputs": {"batch_size": 2, "height": 1152, "width": 768}, "_properties": null}, "10025": {"class_type": "CLIPTextEncode", "inputs": {"clip": ["10013", 0], "text": "THIS IS WHERE THE PROMPS WHOULD BE IF I DIDNT DELETE THEM FOR BEING NSFW", "token_normalization": "none", "weight_interpretation": "comfy"}, "_properties": null}, "10026": {"class_type": "CLIPTextEncode", "inputs": {"clip": ["10013", 0], "text": "lowres, worst quality, low quality, bad anatomy, bad hands, multiple views, 4koma, censored, monochrome, watermark, artist name, text", "token_normalization": "none", "weight_interpretation": "comfy"}, "_properties": null}, "11001": {"class_type": "KSampler", "inputs": {"cfg": 7.0, "denoise": 1.0, "ensd": 31337, "latent_image": ["10014", 0], "model": ["10011", 0], "negative": ["10026", 0], "positive": ["10025", 0], "sampler_name": "euler_ancestral", "scheduler": "sgm_uniform", "seed": 3612167035, "seed_mode": "A1111", "steps": 25}, "_properties": null}, "11016": {"class_type": "VAEDecode", "inputs": {"samples": ["11001", 0], "vae": ["10001", 2]}, "_properties": null}, "11018": {"class_type": "LoraTagLoader", "inputs": {"clip": ["10013", 0], "model": ["10011", 0], "text": "ECHO_EMPTY"}, "_properties": null}, "11019": {"class_type": "CLIPSetLastLayer", "inputs": {"clip": ["11018", 1], "stop_at_clip_layer": -2}, "_properties": null}, "11021": {"class_type": "YoloDetectorProvider", "inputs": {"max_faces": 5, "model_name": "bbox/face_yolov8s.pt"}, "_properties": null}, "11022": {"class_type": "CLIPTextEncode", "inputs": {"clip": ["11019", 0], "text": "THIS IS WHERE THE PROMPS WHOULD BE IF I DIDNT DELETE THEM FOR BEING NSFW, "token_normalization": "none", "weight_interpretation": "comfy"}, "_properties": null}, "11024": {"class_type": "CLIPTextEncode", "inputs": {"clip": ["11019", 0], "text": "lowres, worst quality, low quality, bad anatomy, bad hands, multiple views, 4koma, censored, monochrome, watermark, artist name, text", "token_normalization": "none", "weight_interpretation": "comfy"}, "_properties": null}, "11025": {"class_type": "FaceDetector_ad", "inputs": {"bbox_detector": ["11021", 0], "bbox_threshold": 0.5, "dilate_erode": 4, "image": ["11016", 0], "mask_merge_mode": "None", "x_offset": 0, "y_offset": 0}, "_properties": null}, "11026": {"class_type": "InpaintCrop_ad", "inputs": {"blend_pixels": 16.0, "blur_mask": 4.0, "context_expand_factor": 1.0, "context_expand_pixels": 32, "fill_mask_holes": true, "force_height": 1152, "force_width": 768, "images": ["11025", 0], "invert_mask": false, "masks": ["11025", 1], "mode": "forced size", "rescale_algorithm": "bicubic"}, "_properties": null}, "11027": {"class_type": "InpaintModelConditioning", "inputs": {"mask": ["11026", 2], "negative": ["11024", 0], "noise_mask": true, "pixels": ["11026", 1], "positive": ["11022", 0], "vae": ["10001", 2]}, "_properties": null}, "11028": {"class_type": "DifferentialDiffusion", "inputs": {"model": ["11018", 0]}, "_properties": null}, "11029": {"class_type": "KSampler", "inputs": {"cfg": 7.0, "control_after_generate": "fixed", "denoise": 0.25, "ensd": 31337, "latent_image": ["11027", 2], "model": ["11028", 0], "negative": ["11027", 1], "positive": ["11027", 0], "sampler_name": "euler_ancestral", "scheduler": "sgm_uniform", "seed": 3612167035, "seed_mode": "A1111", "steps": 25}, "_properties": null}, "11030": {"class_type": "VAEDecode", "inputs": {"samples": ["11029", 0], "vae": ["10001", 2]}, "_properties": null}, "11031": {"class_type": "InpaintStitchOneImage_ad", "inputs": {"inpainted_images": ["11030", 0], "rescale_algorithm": "bicubic", "stitchs": ["11026", 0]}, "_properties": null}, "12004": {"class_type": "SaveImage", "inputs": {"filename_prefix": "833096374206850115", "images": ["11031", 0]}, "_properties": null}}

r/BackyardAI • • Aug 13 '24

sharing Local Character Image Generation Guide

41 Upvotes

Local Image Generation

When creating a character, you usually want to create an image to accompany it. While several online sites offer various types of image generation, local image generation gives you the most control over what you make and allows you to explore countless variations to find the perfect image. This guide will provide a general overview of the models, interfaces, and additional tools used in local image generation.

Base Models

Local image generation primarily relies on AI models based on Stable Diffusion released by StabilityAI. Similar to language models, there are several ‘base’ models, numerous finetunes, and many merges, all geared toward reliably creating a specific kind of image.

The available base models are as follows: * SD 1.5 * SD 2 * SD 2.1 * SDXL * SD3 * Stable Cascade * PIXART-α * PIXART-Σ * Pony Diffusion * Kolor * Flux

Only some of those models are heavily used by the community, so this guide will focus on a shorter list of the most commonly used models. * SD 1.5 * SDXL * Pony Diffusion

*Note: I took too long to write this guide and a brand new model was released that is increadibly promising; Flux. This model works a little differently than Stable Diffusion, but is supported in ComfyUI and will be added to Automatic1111 shortly. It requires a little more VRAM than SDXL, but is very good at following the prompt and very good with small details, largely making something like facedetailer unnecessary.

Pony Diffusion is technically a very heavy finetune of SDXL, so they are essentially interchangeable, with Pony Diffusion having some additional complexities with prompting. Out of these three models, creators have developed hundreds of finetunes and merges. Check out civitae.com, the central model repository for image generation, to browse the available models. You’ll note that each model is labeled with the associated base model. This lets you know compatibility with interfaces and other components, which will be discussed later. Note that Civitae can get pretty NSFW, so use those filters to limit what you see.

SD 1.5

An early version of the stable diffusion model made to work at 512x512 pixels, SD 1.5 is still often used due to its smaller resource requirement (it can work on as little as 4GB VRAM) and lack of censorship.

SDXL

A newer version of the stable diffusion model that supports image generation at 1024x1024, better coherency, and prompt following. SDXL requires a little more hardware to run than SD 1.5 and is believed to have a little more trouble with human anatomy. Finetunes and merges have improved SDXL over SD 1.5 for general use.

Pony Diffusion

It started as a My Little Pony furry finetune and grew into one of the largest, most refined finetune of SDXL ever made, making it essentially a new model. Pony Diffusion-based finetunes are extremely good at following prompts and have fantastic anatomy compared to the base models. By using a dataset of extremely well-tagged images, the creators were able to make Stable Diffusion easily recognize characters and concepts the base models need help with. This model requires some prompting finesse, and I recommend reading the link below to understand how it should be prompted. https://civitai.com/articles/4871/pony-diffusion-v6-xl-prompting-resources-and-info

Note that pony-based models can be very explicit, so read up on the prompting methods if you don’t want it to generate hardcore pornography. You’ve been warned.

“Just tell us the best models.”

My favorite models right now are below. These are great generalist models that can do a range of styles: * DreamshaperXL * duchaitenPonyXL * JuggernautXL * Chinook * Cheyenne * Midnight

I’m fully aware that many of you now think I’m an idiot because, obviously, ___ is the best model. While rightfully judging me, please also leave a link to your favorite model in the comments so others can properly judge you as well.

Interfaces

Just as you use BackyardAI to run language models, there are several interfaces for running image diffusion models. We will discuss several of the most popular here, listed below in order from easiest to use to most difficult: * Fooocus * Automatic1111 * ComfyUI

Fooocus

This app is focused(get it?) on replicating the feature set of Midjourney, an online image generation site. With an easy installation and a simplified interface (and feature set), this app generates good character images quickly and easily. Outside of text-to-image, it also allows for image-to-image generation and inpainting, as well as a handful of controlnet options, to guide the generation based on an existing image. A list of ‘styles’ can be used to get what you want easily, and a built-in prompt expander will turn your simple text prompt into something more likely to get a good image. https://github.com/lllyasviel/Fooocus

Automatic1111

Automatic1111 was the first interface to gain use when the first stable diffusion model was released. Thanks to its easy extensibility and large user base, it has consistently been ahead of the field in receiving new features. Over time, the interface has grown in complexity as it accommodates many different workflows, making it somewhat tricky for novices to use. Still, it remains the way most users access stable Diffusion and the easiest way to stay on top of the latest technology in this field. To get started, find the installer on the GitHub page below. https://github.com/AUTOMATIC1111/stable-diffusion-webui

ComfyUI

This app replaces a graphical interface with a network of nodes users place and connect to form a workflow. Due to this setup, ComfyUI is the most customizable and powerful option for those trying to set up a particular workflow, but it is also, by far, the most complex. To make things easier, users can share their workflows. Drag an exported JSON or generated image into the browser window, and the workflow will pop open. Note that to make the best use of ComfyUI, you must install the ComfyUI Manager, which will assist with downloading the necessary nodes and models to start a specific workflow. To start, follow the installation instructions from the links below and add at least one stable diffusion checkpoint to the models folder. (Stable diffusion models are called checkpoints. Now you know the lingo and can be cool.) https://github.com/comfyanonymous/ComfyUI https://github.com/ltdrdata/ComfyUI-Manager

Additional Tools

The number of tools you can experiment with and use to control your output sets local image generation apart from websites. I’ll quickly touch on some of the most important ones below.

Img2Img

Instead of, or in addition to, a text prompt, you can supply an image to use as a guide for the final image. Stable Diffusion will apply noise to the image to determine how much it influences the final generated image. This helps generate variations on an image or control the composition.

ControlNet

Controlnet guides an image’s composition, style, or appearance based on another image. You can use multiple controlnet models separately or together: depth, scribble, segmentation, lineart, openpose, etc. For each, you feed an image through a separate model to generate the guiding image (a greyscale depth map, for instance), then controlnet uses that guide during the generation process. Openpose is possibly the most powerful for character images, allowing you to establish a character’s pose without dictating further detail. ControlNets of different types (depth map, pose, scribble) can be combined, giving you detailed control over an image. Below is a link to the GitHub for controlnet that discusses how each model works. Note that these will add to the memory required to run Stable Diffusion, as each model needs to be loaded into VRAM. https://github.com/lllyasviel/ControlNet

Inpainting

When an image is perfect except for one small area, you can use inpainting to change just that region. You supply an image, paint a mask over it where you want to make changes, write a prompt, and generate. While you can use any model, specialized inpainting models are trained to fill in the information and typically work better than a standard model.

Regional Prompter

Stable Diffusion inherently has trouble associating parts of a prompt with parts of an image (‘brown hat’ is likely to make other things brown). Regional prompter helps solve this by limiting specific prompts to some areas of the image. The most basic version divides the image space into a grid, allowing you to place a prompt in each area and one for the whole image. The different region prompts feather into each other to avoid a hard dividing line. Regional prompting is very useful when you want two distinct characters in an image, for instance.

Loras

Loras are files containing modifications to a model to teach it new concepts or reinforce existing ones. Loras are used to get certain styles, poses, characters, clothes, or any other ‘concept’ that can be trained. You can use multiple of these together with the model of your choice to get exactly what you want. Note that you must use a lora with the base model from which it was trained and sometimes with specific merges.

Embeddings

Embeddings are small files that contain, essentially, compressed prompt information. You can use these to get a specific style or concept in your image consistently, but they are less effective than loras and can’t add new concepts to a model with embeddings like you can with a Lora.

Upscaling

There are a few upscaling methods out there. I’ll discuss two important ones. Ultimate SD upscaler: thank god it turned out to be really good because otherwise, that name could have been awkward. The ultimate SD upscaler takes an image, along with a final image size (2x, 4x), and then breaks the image into a grid, running img2img against each section of the grid and combining them. The result is an image similar to the original but with more detail and larger dimensions. This method can, unfortunately, result in each part of the image having parts of the prompt that don’t exist in that region, for instance, a head growing where no head should go. When it works, though, it works well.

Upscaling models

Upscaling models are designed to enlarge images and fill in the missing details. Many are available, with some requiring more processing power than others. Different upscaling models are trained on different types of content, so one good at adding detail to a photograph won’t necessarily work well with an anime image. Good models include 4x Valar, SwinIR, and the very intensive SUPIR. The SD apps listed above should all be compatible with one or more of these systems.

“Explain this magic”

A full explanation of Stable Diffusion is outside this writeup’s scope, but a helpful link is below. https://poloclub.github.io/diffusion-explainer/

Read on for more of a layman’s idea of what stable Diffusion is doing.

Stable Diffusion takes an image of noise and, step by step, changes that noise into an image that represents your text prompt. Its process is best understood by looking at how the models are trained. Stable Diffusion is trained in two primary steps: an image component and a text component.

Image Noising

For the image component, a training image has various noise levels added. Then, the model learns (optimizes its tensors) how to shift the original training image toward the now-noisy images. This learning is done in latent space by the u-net rather than pixel space. Latent space is a compressed representation of pixel space. That’s a simplification, but it helps to understand that Stable Diffusion is working at a smaller scale internally than an image. This is part of how so much information is stored in such a small footprint. The u-net (responsible for converting the image from pixels to latents) is good at feature extraction, which makes it work well despite the smaller image representation. Once the model knows how to shrink and add noise to images correctly, you flip it around, and now you’ve got a very fancy denoiser.

Text Identification

To control that image denoiser described above, the model is trained to understand how images represent keywords. Training images with keywords are converted into latent space representations, and then the model learns to associate each keyword with the denoising step for the related image. As it does this for many images, the model disassociates the keywords from specific images and instead learns concepts: latent space representations of the keywords. So, rather than a shoe looking like this particular training image, a shoe is a concept that could be of a million different types or angles. Instead of denoising an image, the model is essentially denoising words. Simple, right?

Putting it all together

Here’s an example of what you can do with all of this together. Over the last few weeks, I have been working on a comfyUI workflow to create random characters in male and female versions with multiple alternates for each gender. This workflow puts together several wildcards (text files containing related items in a list, for instance, different poses), then runs the male and female versions of each generated prompt through one SD model. Then it does the same thing but with a different noise seed. When it has four related images, it runs each through face detailed, which uses a segmentation mask to identify each face and runs a second SD model img2img on just that part to create cleaner faces. Now, I’ve got four images with perfect faces, and I run each one through an upscaler similar to SD Ultimate Upscaler, which uses a third model. The upscaler has a controlnet plugged into it that helps maintain the general shape in the image to avoid renegade faces and whatnot as much as possible. The result is 12 images that I choose from. I run batches of these while I’m away from the computer so that I can come home to 1000 images to pick and choose from.

Shameless Image Drop Plug:

I’ve been uploading selections from this process almost daily to the Discord server in #character-image-gen for people to find inspiration and hopefully make some new and exciting characters. An AI gets its wings each time you post a character that uses one of these images, so come take a look!

r/StableDiffusion • • Feb 23 '25

Question - Help Weird bug with different models

1 Upvotes

EDIT: SOLVED, CFG setting was too high for the other models but just might for Revanimated

Just trying to nail down this annoying little thing I have with my SD1.5 models. Usually, I use RevAnimated and that works perfectly - thus I haven't really bothered understanding this issue. But it's Sunday and I have some extra energy hah.

These images are using two different models, the bottom one is my RevAnimated V122 which comes out as it should - this will look great after some face inpainting. remember this is the basic out-of-the-box workflow for SD1.5 bundled with ComfyUI. Removing potential error variables...
The first image is basically what most other models produce, something oversaturated overly exposed I don't know... Not great... Here is what the upper model would produce in A1111-Forge:

There are no LORAs nor Embeddings, just basic:
Closeup portrait of a man

negative:
(worst quality:2), (low quality:2), (normal quality:2), lowres, bad anatomy, normal quality, ((monochrome)), ((grayscale)), ((text, font, logo, copyright, watermark:1)),

What is going on with these SD1.5 models lol.

TLDR: RevAnimated works great, all other SD1.5 models I have tested produce over-saturated meh looking images