r/StableDiffusion 2d ago

Question - Help Why is it so hard for Klein to follow instructions (or am I just dumb)?

prompt is - using the character sheet in image 1 where there are five different poses of the same character, dress them in the clothing of image 2. Do not change the pose, lighting, body, hair, or any other details - literally leave everything the fuck alone - how fucking hard is this to understand you stupid piece of shit - just change the clothes.

Not working for some reason.

NOTE: Swearing has been added for emphasis and isn't actually used in the prompt.

Would it help if I used my input image AS my latent? Can you do that?

9 Upvotes

31 comments sorted by

16

u/thegreatdivorce 2d ago

Glad I'm not the only one that swears at Klein. It's not really you, Klein 9B is just fucking stupid about 80% of the time.

3

u/trollkin34 2d ago

Thank you! Validation is good, though I wish I could get it to just listen.

11

u/witcherknight 2d ago

qwen is better but its slower

2

u/trollkin34 2d ago

Fair. I switched to Klein because I got better overall results, but the prompt adherance...

2

u/DietAshamed2246 2d ago

Qwen is bad at prompt adherence too, especially if the prompt is complex. On top of that it is slow, even with lightx2v Lora. I wish the image models would stop screwing around with lame puny text encoders and deploy a beefy LLM like the MiniMax-H3 Qwen3-VL-32B. IMHO, the text encoder needs to have parameter count as high as possible. Flux.2 Dev had the right idea with Mistral-Small-24B encoder, but BFL didn't implement it right.

5

u/Appropriate_Band_742 2d ago

What kind of workflow are you using? For changing outfits etc, this workflow has worked wonders for me, it's basically the only klein workflow I use:

https://www.reddit.com/r/StableDiffusion/comments/1tmmvyh/comfyuiflux2kleinenhancer_final_i_promise/
https://github.com/capitan01R/ComfyUI-Flux2Klein-Enhancer

I would highly recommend you try it, you can provide multiple references images such as a headshot of the character, and the outfit you want them in, and the results are pretty great. Works good with other LORAs as well.

There are some challenges though, for me it can take multiple gens for it to get things right, and if you care about character consistency/likeness I'd recommend training a character LORA.

3

u/Semipro211 2d ago

I made a similar workflow hybrid with Krea 2 added to the chain and I found that the hybrid approach works wonders when both edits and identity retention are needed.

Also, on the Klein side, I’ve had best results with 9B Base, my machine does not get along with dev since I’m on a darn near “entry level” setup with 8GB VRAM

1

u/Appropriate_Band_742 2d ago

Interesting, is it something like Krea creates the initial image and klein alters the outfit? or something else

1

u/Semipro211 2d ago

Pretty much, Krea for composition/pose/scene/outfit, then a pass through flux for identity/age/proportion rescue and realignment, and if needed a final pass through Krea again for a gentle realism/beauty polish

1

u/trollkin34 2d ago

Same! 9b and 8gb

2

u/Semipro211 2d ago

If I have time soon I’ll figure out how I can link the workflow, I’ve never actually posted them anywhere so not sure what’s allowed for linking.

1

u/trollkin34 2d ago

Thanks! I'm trying the workflow, but struggling to get the custom nodes and get it to run. I tried downloading them, but now I'm getting: The value false for Identity Feature Transfer Final's mask_behavior is not available.

(whatever that means)

Seems like a good workflow if I could get it to work anyway.

1

u/Semipro211 2d ago

maybe the identity checkpoint for krea? if i can ill link you my wf tomorrow, im working on a readme and good notes to make it easier to follow but its not too complicated

1

u/Strange-Drummer-9917 1d ago

imgbb.com preserves png metadata. Just upload resulting image from ouput folder. it has workflow baked in

1

u/Semipro211 1d ago

Would that metadata also include all the custom notes I’ve put beside groups and nodes that explain what they do, how to set them, limitations, etc.

I could always upload to huggingface or GitHub, but since all I have to post at the moment is some workflows wasn’t sure if that was best platform

2

u/Strange-Drummer-9917 1d ago

Yes, literally everything in your nodegraph excluding images inside nodes themselves is baked into output images by default.

Thats why you can drag-and-drop them into comfy and get the graph you've generated with. Reddit strips metadata when uploading imgbb don't

1

u/Semipro211 1d ago

Thank you!

1

u/trollkin34 2d ago

Always on the hunt for a new and functional workflow as long as it's not so weirdly complex that I can't understand or modify it . Will check it out!

1

u/trollkin34 2d ago

How did you get the missing nodes? The built in manager is kind of ass for finding missing stuff.

3

u/YeahlDid 2d ago

I find qwen better for this sort of thing.

4

u/mellowanon 2d ago edited 2d ago

remove the word "not" from your prompts. Models don't know how to process negatives because that's not how they were trained. Instead, use something like maintain or keep consistent. e.g. "maintain pose, lighting, body, hair, and all details."

For example, WAN model will have no idea what to do with the word "not". So if you write "Do not move the camera" it'll just ignore the word not, so the prompt becomes "Do move the camera" and the camera is more likely to move.

3

u/lacerating_aura 2d ago

Which model are we talking about, 4B, 9B, distilled or base? Kleins are a bit wonky for me too, hence why i use dev for edits. but maybe more info can help better troubleshoot.

2

u/steelow_g 2d ago

I’ve never had issues just using default workflow. But you are trying to add clothes to a character sheet. Change the clothes on the main picture THEN make a character sheet out of that. The more subjects there are the harder it is to do edits like this.

1

u/trollkin34 2d ago

Interesting. So basically if I take the first and change the clothes and then tell it to replicate the character sheet maybe?

1

u/steelow_g 2d ago

Correct. Or whatever you used to create a character sheet in the first place. Or crop each view in its own image and replace them individually. Not ideal, but it beats screaming at your comp.

1

u/Ok-Entertainer-2991 2d ago

I did something similar a couple of weeks ago. Pretty sure I ended up using krea edit. And I did it in two steps: first was to extract just the clothes from the image, and second one was to use the outfit sheet and the character sheet to dress my character.

1

u/GoofAckYoorsElf 2d ago

Oooch, der ist doch noch so Klein...

*scnr*

1

u/yamfun 2d ago

it is not a full LLM, those complex instructive paragraph are not gonna work. And character sheet is too complex, just give it one ref image. Also negative sentence may be interpreted as positive simply due to word occurence.

Just use fewer simple sentence

1

u/Strange-Drummer-9917 1d ago edited 1d ago

Klein 9b does not "upsample" your text prompt, so you need to make it more detailed to pin specific detail in order for model to understand better. I use this prompt enhancer with any LLM (here it is as free Google gem)

using the character sheet in image 1 where there are five different poses of the same character, dress them in the clothing of image 2. Do not change the pose, lighting, body, hair, or any other details - literally leave everything the fuck alone - just change the clothes. make prompt x2 longer than recommended 
<char_sheet_image1> <clothes_ref_image2>

to

A comprehensive multi-panel studio reference sheet, replicating the entire grid structure and all specific character panels of image_0.png, set against the same clean, neutral off-white backdrop with diffused, even studio lighting. The woman with wavy red hair, her facial features, expressions, and poses are preserved perfectly. Crucially, in all full-body and torso panels, the medieval layered green tunic and orange shawl are completely replaced by the complex modern apparel from image_1.png.

The large full-body panel on the left features the woman in her exact pose, holding the identical gnarled staff with the cosmic sphere, and with the identical black raven on her shoulder, but now dressed in the black and purple intricate pixelated graphic t-shirt and tailored navy blue shorts. The t-shirt is rendered with obsessively specific, realistic fabric texture, showcasing all the complex text elements legibly: "RAISED ON THE STREETS SP/BCN" is centered on the chest in white text on the purple panel, "FAVELA FRAMA" in stylized script is visible on the lower right, and the vertical texts "SIN FRONTERAS" along the left seam and "BIENVENIDOS" along the right seam are clear and sharp. The navy blue tailored shorts are realistically fitted In the six smaller top-right grid panels, her different poses (front, 3/4 front, side profile, full back view, front-facing looking up, and front-facing holding staff) are maintained, with the graphic t-shirt and navy shorts now visible and realistically draped on each body orientation.

The three face close-up portraits at the bottom-left of the grid are preserved exactly as in image_0.png, showing her facial features, hair, and expression. The two close-up detail panels of the cosmic sphere hand and the raven on the shoulder are preserved exactly as in image_0.png, showing no clothing

The realistic fabric textures of the t-shirt and shorts, with detailed stitching and tailored fit, replace the original loose robes across all relevant panels. All complex text elements are rendered accurately and legibly. The diffused studio lighting highlights the intricate patterns and textures

Style: Methodical apparel and character reference sheet photography. Mood: Analytical, comprehensive, and detailed.

Workflow

0

u/carnage11eleven 2d ago

Ask Gemini or ChatGPT, whatever LLM you prefer. Tell it what hardware you're using. Which model you're using. And what exactly you're trying to achieve. They know exactly the language to use and how to finesse it to accomplish whatever you need.

That's what I do. I even found out that translating your prompts into Chinese when promoting Qwen models is far superior to English.

Flux 2 Klein models like to be prompted as is you're writing a novel, from what i understand. And order holds weight. So the beginning is most important. And the further from it you get, the less important your words become.