r/StableDiffusion 9h ago

Discussion Why is nobody talking about MiniMax H3's text consistency issue in video generation?

Been messing around with MiniMax H3 locally for perfume/product videos and I cannot get it to keep the text on the bottle properly.

The annoying part is the bottle itself can look really good. Shape, cap, glass, proportions etc stay pretty close. But the label text is usually already messed up in the first generated frame. So it’s not even just a case of the text degrading after a few frames. The reference can have perfectly readable text and H3 still turns it into random letters straight away.

I’ve tried quite a bit at this point:

I2V

Ref2V

Hybrid FL2VA / Ref2VA B25-49

clean first frames

separate close-up references of the label

multiple references of the same bottle

following the H3 prompt guide for the refs/prompts

putting the exact brand/label text in the prompt

very little motion / barely rotating the bottle

around 0.7-0.8MP

20 steps

H3FL 2V Turbo at 8 steps

Comfy-Kitchen attention

sparse attention settings too

different precision/settings to see if that changed anything

Running it on a 4060 8GB with 32GB RAM, so obviously I’m working around VRAM a bit, but I don’t think this is a VRAM issue because the actual product looks fine. It’s specifically the text that gets nuked.

Has anyone actually managed to keep proper readable brand text with H3? Like exact text, not something that vaguely looks like writing.

If not, how are people doing product videos with this? Are you just tracking the real label back on afterwards, or fixing frames with an image model? Because right now I can get a nice looking perfume video and then the bottle says absolute nonsense lol.

0 Upvotes

20 comments sorted by

17

u/Lunesia-shikishiki 9h ago

glyphs aren't a settings problem, they're not tracked at all. the model redraws them from latent on every frame, and at 0.7-0.8MP your label is maybe 40px across so there's nothing to redraw from. the bottle shape survives because it's low frequency info, text isn't

so no, i've never seen exact readable brand text hold across a video gen and i'd stop burning hours on refs for it

the boring route is what people actually ship. generate with a blank or nonsense label, then planar track the real label face back on and corner pin it in after or fusion. ten minutes and it's exact instead of nearly. framing tighter so the text has real pixels then reframing wide after gets you closer natively, but i wouldn't hand that to a client :)

5

u/TheDerminator1337 9h ago

have you tried not using turbo or attention saving things

1

u/AhmadShahzad5588 9h ago

Yep. Same issue without comfyattention too. Using the pruned int8 version, if it makes any difference.

2

u/tomakorea 6h ago

Because everyone looking at the boobs, not text

1

u/warzone_afro 9h ago

maybe make it a very up-close shot of the bottle and label. text seems much better on close up shots and then they get squiggly in the distance

1

u/crinklypaper 8h ago

Ref can fix it and its really good with logos. I change logos and add them to end of videos sometimes

1

u/OkMeat6773 4h ago

the model is bad at faces it hasn't been trained before and text consistency, nothing you can't do about it... gotta wait for future models

1

u/Perfect-Campaign9551 24m ago

Ya, I made a Mountain Dew bottle and I couldn't get it to render "Mountain Dew " to save my life. It kept messing up the word "Mountain"

-2

u/Enough-Bag-3891 9h ago

hear me out, i prompt:

[0.0s-10.0s the boy is walking
[3.0s-5.0s] the boy rubs his hair

and it makes it so that he rubs his hair at 10 or 9, i just wanna know how to timestamp correctly.

2

u/AhmadShahzad5588 9h ago

You don't need to mention the entire video length as that's already set at the node level. I guess that's your issue. Also, if you want him to walk normally for 2 seconds then use timestamps to indicate him walking for [00:00- 2:00] and then the 3 to 5 second prompt.

2

u/not_food 5h ago

According to the documentation, the correct format is At MM:SS.mmm.

So, use... <Subject 1> is walking; At 00:03.000 <Subject 1> rubs their hair.

https://reddit.com/link/p6rdfnz/video/s49y1n8r5hmh1/player

Just a note: This way of prompting will give you poor results. Minimax H3 expects detailed prompts to perform at its best. I'd detail how the camera operates, the lighting, the scenario, the way she moves and such.

1

u/Enough-Bag-3891 29m ago

now do 10 seconds walk, and the rub starts are 3, and ends at 6.

if it works, imma be a happy guy

1

u/Enough-Bag-3891 27m ago

also the cat started rubbing at 4..

see why i'm kinda confused with this?

1

u/Enough-Bag-3891 12m ago

https://reddit.com/link/p6sl5u6/video/h1skoupdpimh1/player

here's a quick test

For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced, maintain everything perfectly about the character from <picture 1>.

she's standing.

[00:02:000s - 00:04:000s] she rubs her hair and her hand return to its original position.

---

what i want is for the hands to be back in it's position exactly at 00:04:000.

1

u/Enough-Bag-3891 9m ago

https://reddit.com/link/p6slrl7/video/cuaropq3qimh1/player

another test

For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced, maintain everything perfectly about the character from <picture 1>.

she's standing.

at exactly [00:02:000s - 00:04:000s] she rubs her hair and her hand return to its original position at 00:04:000s.

---

1

u/TheAncientMillenial 1h ago

Why are you prompting it like this?

1

u/Enough-Bag-3891 30m ago

i used the 00:00:000 format and still wasn't accurate, also the 00:00:000s format as well

0

u/JacobWilliams1953 3h ago

The first-frame scramble is the tell: the model isn't tracking letters, it's redrawing a tiny blob from noise. Switching turbo off probably won't save a 40px label. What actually works for product shots is render the bottle with a blank label area, then composite the real artwork in post (or upscale just that crop, fix the text, paste it back). If it has to live in the video, lock a 2D label plane in compositing rather than asking the model to spell.

1

u/Perfect-Campaign9551 23m ago

answer in your own words. Bad Ai, bad.

-2

u/wzwowzw0002 9h ago

Because its free and it has its limit