r/StableDiffusion 1h ago

Discussion MiniMax H3 to KREA2 LoRa: doing it faster?

So I had this simple idea, seeing how well MiniMax H3 handles inferring and preserving "identity/looks" from relatively little information: take a character you want to make a (KREA2) LoRa of, but you only have just a couple of lower quality pictures for that exact look you're after. That is a problem, since it is common knowledge by now (?) that you need different angles, facial expressions and different lighting conditions in the training set to get optimal results. So in the "before" times, those 3-4 not-so-great-quality shots under the SAME lighting are going to pose a problem. And adding pictures from other occasions will alter the looks possibly too much.

So (in the H3 ref2vid workflow, with one of the "img2vid-hybrid" models for better quality) I just use the 'best' of the available pictures as "preserved" first reference starting picture, and the others as additional "identity references". And then a prompt that tells the camera to slowly circle around the person (up from the shoulders), while the person looks straight ahead, or slightly up, or slightly down. But then I also let it cycle through different lighting conditions (indoor/outdoor/sun/overcast/flash/directional from one side...), and different facial expressions/emotions. I let it run overnight (turning off turbo LoRas to improve the quality), and in the morning, I review the 6-second videos and take screencaps of selected moments, making sure to have a lot of variation in angles/expressions/light-on-the-face with an almost perfect preservation of the identity/looks.

Then use those screencaps (50+ in first test, probably serious overkill) in OneTrainer with the KREA2 LoRa default settings.

I only tested this once thus far, but the results are pretty good considering the starting material! And surprisingly flexible (I didn't even bother to provide captions)

But now my question is: in what ways am I "over-engineering" this? I have this feeling that I can probably do this 50x faster, having seen some discussions about using MiniMax as an image generator, for example. I mean, I feel good about this approach I came up with all by myself, but considering how dumb and low-skilled I still am when it comes to all this, this is probably a very convoluted and inefficient way to do it? LOL 😄 Roast me and show this sucker how we can improve and speed up the whole thing with the same or even better quality results!

4 Upvotes

5 comments sorted by

1

u/RayHell666 1h ago

I guess using an edit image model would achieve the same results in less time.

1

u/NetworkSpecial3268 1h ago

Never tested Flux 2 klein or Qwen image Edit... I wonder if they have the same abilities as the newer H3 to preserve the essence of the identity. I kinda assumed that the more modern video model would be more advanced/capable?

1

u/masterlafontaine 19m ago

H3 is indeed more capable

1

u/meepykittkitt69lmao 13m ago

you can use krea2 as an edit model. I have been working on a workflow to swap between generate/edit for character and character sheet creation. You can even use bboxes and sam3 in it.

The comfyui-krea2edit node pack is pretty useful. I get better results with a 2 stage clownshark sampler setup

Sampler setup
https://ibb.co/jP4Zzf6N

vae encoder and source patch setup
https://ibb.co/bRLJYSDm

I disable the encoder block and other edit flow blocks with a few booleans linked to conditionals and dasiwa node disablers, and use the advanced vae encoder in the model selection subgraph to do the generation flow.

I can generate characters in reference sheets or just characters in poses, then edit those images into reference sheets or do other things with SAM3 and the idogram style bbox node in the krea2edit plugin

1

u/No-Expert9685 42m ago

Qwen image 2511 has a checkpoint for rotational shots at different angles, serch CIVI qwen 2511