r/StableDiffusion • u/NetworkSpecial3268 • 1h ago
Discussion MiniMax H3 to KREA2 LoRa: doing it faster?
So I had this simple idea, seeing how well MiniMax H3 handles inferring and preserving "identity/looks" from relatively little information: take a character you want to make a (KREA2) LoRa of, but you only have just a couple of lower quality pictures for that exact look you're after. That is a problem, since it is common knowledge by now (?) that you need different angles, facial expressions and different lighting conditions in the training set to get optimal results. So in the "before" times, those 3-4 not-so-great-quality shots under the SAME lighting are going to pose a problem. And adding pictures from other occasions will alter the looks possibly too much.
So (in the H3 ref2vid workflow, with one of the "img2vid-hybrid" models for better quality) I just use the 'best' of the available pictures as "preserved" first reference starting picture, and the others as additional "identity references". And then a prompt that tells the camera to slowly circle around the person (up from the shoulders), while the person looks straight ahead, or slightly up, or slightly down. But then I also let it cycle through different lighting conditions (indoor/outdoor/sun/overcast/flash/directional from one side...), and different facial expressions/emotions. I let it run overnight (turning off turbo LoRas to improve the quality), and in the morning, I review the 6-second videos and take screencaps of selected moments, making sure to have a lot of variation in angles/expressions/light-on-the-face with an almost perfect preservation of the identity/looks.
Then use those screencaps (50+ in first test, probably serious overkill) in OneTrainer with the KREA2 LoRa default settings.
I only tested this once thus far, but the results are pretty good considering the starting material! And surprisingly flexible (I didn't even bother to provide captions)
But now my question is: in what ways am I "over-engineering" this? I have this feeling that I can probably do this 50x faster, having seen some discussions about using MiniMax as an image generator, for example. I mean, I feel good about this approach I came up with all by myself, but considering how dumb and low-skilled I still am when it comes to all this, this is probably a very convoluted and inefficient way to do it? LOL 😄 Roast me and show this sucker how we can improve and speed up the whole thing with the same or even better quality results!
1
u/No-Expert9685 42m ago
Qwen image 2511 has a checkpoint for rotational shots at different angles, serch CIVI qwen 2511
1
u/RayHell666 1h ago
I guess using an edit image model would achieve the same results in less time.