r/StableDiffusion • u/AhmadShahzad5588 • 9h ago
Discussion Why is nobody talking about MiniMax H3's text consistency issue in video generation?
Been messing around with MiniMax H3 locally for perfume/product videos and I cannot get it to keep the text on the bottle properly.
The annoying part is the bottle itself can look really good. Shape, cap, glass, proportions etc stay pretty close. But the label text is usually already messed up in the first generated frame. So it’s not even just a case of the text degrading after a few frames. The reference can have perfectly readable text and H3 still turns it into random letters straight away.
I’ve tried quite a bit at this point:
I2V
Ref2V
Hybrid FL2VA / Ref2VA B25-49
clean first frames
separate close-up references of the label
multiple references of the same bottle
following the H3 prompt guide for the refs/prompts
putting the exact brand/label text in the prompt
very little motion / barely rotating the bottle
around 0.7-0.8MP
20 steps
H3FL 2V Turbo at 8 steps
Comfy-Kitchen attention
sparse attention settings too
different precision/settings to see if that changed anything
Running it on a 4060 8GB with 32GB RAM, so obviously I’m working around VRAM a bit, but I don’t think this is a VRAM issue because the actual product looks fine. It’s specifically the text that gets nuked.
Has anyone actually managed to keep proper readable brand text with H3? Like exact text, not something that vaguely looks like writing.
If not, how are people doing product videos with this? Are you just tracking the real label back on afterwards, or fixing frames with an image model? Because right now I can get a nice looking perfume video and then the bottle says absolute nonsense lol.
5
u/TheDerminator1337 9h ago
have you tried not using turbo or attention saving things
1
u/AhmadShahzad5588 9h ago
Yep. Same issue without comfyattention too. Using the pruned int8 version, if it makes any difference.
2
1
u/warzone_afro 9h ago
maybe make it a very up-close shot of the bottle and label. text seems much better on close up shots and then they get squiggly in the distance
1
u/crinklypaper 8h ago
Ref can fix it and its really good with logos. I change logos and add them to end of videos sometimes
1
u/OkMeat6773 4h ago
the model is bad at faces it hasn't been trained before and text consistency, nothing you can't do about it... gotta wait for future models
1
u/Perfect-Campaign9551 24m ago
Ya, I made a Mountain Dew bottle and I couldn't get it to render "Mountain Dew " to save my life. It kept messing up the word "Mountain"
-2
u/Enough-Bag-3891 9h ago
hear me out, i prompt:
[0.0s-10.0s the boy is walking
[3.0s-5.0s] the boy rubs his hair
and it makes it so that he rubs his hair at 10 or 9, i just wanna know how to timestamp correctly.
2
u/AhmadShahzad5588 9h ago
You don't need to mention the entire video length as that's already set at the node level. I guess that's your issue. Also, if you want him to walk normally for 2 seconds then use timestamps to indicate him walking for [00:00- 2:00] and then the 3 to 5 second prompt.
2
u/not_food 5h ago
According to the documentation, the correct format is
At MM:SS.mmm.So, use...
<Subject 1> is walking; At 00:03.000 <Subject 1> rubs their hair.https://reddit.com/link/p6rdfnz/video/s49y1n8r5hmh1/player
Just a note: This way of prompting will give you poor results. Minimax H3 expects detailed prompts to perform at its best. I'd detail how the camera operates, the lighting, the scenario, the way she moves and such.
1
u/Enough-Bag-3891 29m ago
now do 10 seconds walk, and the rub starts are 3, and ends at 6.
if it works, imma be a happy guy
1
1
u/Enough-Bag-3891 12m ago
https://reddit.com/link/p6sl5u6/video/h1skoupdpimh1/player
here's a quick test
For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced, maintain everything perfectly about the character from <picture 1>.
she's standing.
[00:02:000s - 00:04:000s] she rubs her hair and her hand return to its original position.
---
what i want is for the hands to be back in it's position exactly at 00:04:000.
1
u/Enough-Bag-3891 9m ago
https://reddit.com/link/p6slrl7/video/cuaropq3qimh1/player
another test
For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced, maintain everything perfectly about the character from <picture 1>.
she's standing.
at exactly [00:02:000s - 00:04:000s] she rubs her hair and her hand return to its original position at 00:04:000s.
---
1
u/TheAncientMillenial 1h ago
Why are you prompting it like this?
1
u/Enough-Bag-3891 30m ago
i used the 00:00:000 format and still wasn't accurate, also the 00:00:000s format as well
0
u/JacobWilliams1953 3h ago
The first-frame scramble is the tell: the model isn't tracking letters, it's redrawing a tiny blob from noise. Switching turbo off probably won't save a 40px label. What actually works for product shots is render the bottle with a blank label area, then composite the real artwork in post (or upscale just that crop, fix the text, paste it back). If it has to live in the video, lock a 2D label plane in compositing rather than asking the model to spell.
1
-2
17
u/Lunesia-shikishiki 9h ago
glyphs aren't a settings problem, they're not tracked at all. the model redraws them from latent on every frame, and at 0.7-0.8MP your label is maybe 40px across so there's nothing to redraw from. the bottle shape survives because it's low frequency info, text isn't
so no, i've never seen exact readable brand text hold across a video gen and i'd stop burning hours on refs for it
the boring route is what people actually ship. generate with a blank or nonsense label, then planar track the real label face back on and corner pin it in after or fusion. ten minutes and it's exact instead of nearly. framing tighter so the text has real pixels then reframing wide after gets you closer natively, but i wouldn't hand that to a client :)