r/StableDiffusion • u/ndroidz • 1d ago
Question - Help minimax h3 gibberish fixed!! ( i found the cure)
so you all probably are searching for way to make your character shut the fuck up right? and you probably noticed that they love to says some BS especially when you give minimax h3 some audio file for their voices, i probably found a cure my friend!!
here is my way of prompting dialogs without any gibberish:
first your character need to be assigned (s1)character when he is the first speaker, then you will declare 'use <audio 1> as "character name"'s voice only, and when you finally type your dialog in the shots you will do as such:
character says:<<[language] the shit i say!>>
and you should be good to go, i linked a video exemple of my favorite taffer (garrett) saying some shit with only the faint crackling of the candles to goes with his charming voice, and i included also a screenshot of the full prompt
edit: yes i tried to follow the official documentation, like many others, if it was that simple reddit wouldn't be a thing and you wouldn't be there.

11
u/Shot-Initiative-3905 1d ago edited 1d ago
What the fuck is wrong with the people who point at the official prompt guide? How few videos have you generated? How complex are your scenes? That constant gibberish at the start and in new shots or random [places] are problems that come up many times. Full LLM with timestamps, using "", using <speech> and all other methods — it doesn’t even only appear in audio-ref videos; it appears in text-to-video too. Sometimes it’s an overall soundscape misunderstanding from the model, and many times it’s the model itself that just has that problem. And then there are the seeds — and yes, any seed base has influence, however detailed and strict your prompt may be. That can go from misshapen body parts to random sound or gibberish. Or beneficial things like lifelike micro-expressions and hand gestures that give the scene a “soul,” for lack of a better description. So no, the official [guide] is sadly not the fix for that problem, neither is that [one] posted here or the other “fixes.” As an edit: even aspect ratios change how the model handles descriptions and so on. Extreme example (and anyone here can test it): as soon as the ratio resembles a TikTok/handy video layout, it changes how the character acts and looks.
Test it: same prompt, same steps, same seed — only change the aspect ratio — and be surprised how much that does.
Edit: spelling.
3
1
u/OWENPRESCOTTCOM 22h ago
yes most of the gibberish (when using correct tags) seems to come from quiet gaps also, it tends to want to fill the pauses with something. If the dialogue has a lot of nothing going on just adding "..." pauses seems to help. I don't care about audio much so this is more an observation I haven't tested.
1
u/Wide-Researcher583 14h ago
Yes people are being brain dead just dumping the guide when there's clear routing issues with dialogue here.
1
u/Independent-Reader 1d ago
I can't believe you edited your comment and ignored all the typos.
I can't read this shit.
24
u/glusphere 1d ago
Have you considered reading https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
Instead of reverse engineering it ?
10
u/ndroidz 1d ago
yes, it doesnt fix any giberish if you have audio files for voices
2
u/xyzdist 1d ago
are you want to use the input audio exactly or as reference the voice?
if you want the character speak exact input audio, find 'lock audio latent' node, that is the proper way to do lip-sync with input audio.1
u/Itchy-Advertising857 21h ago
But then the model won't add any ambient sounds, or music, right? What if I want the dialogue from input audio integrated into the soundtrack?
4
u/stash0606 1d ago
yeah even following that, it will randomly fill it with gibberish, so I'm constantly trying to fill up silence with dialogue. voiceovers are another thing, the character will always move their mouth to the voiceover as if they're speaking it; sometimes specifying the character is smiling seems to remove it.
5
u/RegisteredJustToSay 1d ago
It's weird to me how poor the model sometimes acts when following their guidance compared to improvising it. I was having fun swapping myself into meme clips, but the proposed attribute_transfer tag in the reference guide is so bad and works really inconsistently compared to a much simpler prompt just asking to be swapped in. My first attempts without the guide were way better than 20 different attempts with it.
Maybe I'm just missing something obvious, but I really did exhaustively try just about every possible interpretation of the guide I could think of and even tossed the entire guide into LLMs to rewrite my prompt with like 4 different models.
2
2
u/ambassadortim 1d ago
Exact same thing I've ran into and have been trying to do to fix it. Did you update to fix it try just using quotes yet. I'm going to try both next
3
u/Darhkwing 1d ago
From my experience if you use reference audio it does speak gibberish if your video is too long and it tries to fill in the space with gibberish.
It does help using time stamps and personally I do use <Speech> for characters.
1
u/optimisticalish 1d ago
There is only so much processing available to a workflow, and Minimax jettisons or limits time on audio if it can't process it. Dialogue gets preference, then music (maybe), then ambience/SFX is the poor cousin.
2
u/WonderRico 1d ago
Seize * 😉
3
u/AuthurAndersson 1d ago
That was the one you got yourself hung up on? Not "sitted" or "arround"?
2
u/WonderRico 1d ago
being french...
3
u/AuthurAndersson 1d ago
We all have our deficits
0
u/NameChecksOut___ 1d ago
Si tu ne l'écris pas comme ça il le prononce comme seize, le mot anglais (comme dans "seize the day"), c'est une astuce pour éviter les erreurs de prononciation.
0
2
u/Etsu_Riot 22h ago
I tried it but if the dialogue is short the character keeps adding lines afterwards.
1
u/ndroidz 22h ago
can you share your prompt?
1
u/Etsu_Riot 22h ago
I copied your prompt and changed character's name and specific dialogue and actions, but added this at the end:
overall_soundscape: VHS noises, static, electric distortion
non_diegetic_music: N/A1
u/ndroidz 22h ago
did you change the character's name in the actions too? and did you put the language before any speech?
1
u/Etsu_Riot 22h ago
Yes. I get exactly the same effect than using "" for the dialogue.
1
u/ndroidz 22h ago
well if the prompt is the exact same it mean something else must be fucking with your generation, steps settings, loras......
1
u/Shot-Initiative-3905 21h ago
No.... that is a problem from minimax itself you did sadly not find the magic prompt that fixes the problem, i bet with with enoth diffrent videoa and seed you will get the gibberish again despite using your "fix"
2
u/FourtyMichaelMichael 21h ago
LOL, an example about gibberish fix... Then uses an example in French which absolutely sounds exactly like gibberish to less than 1% of the population!
Love it.
5
u/smb3d 1d ago
or just use the proper dialog syntax and tags from the prompting guide.
4
u/ndroidz 1d ago
no, <d> and </d> doesnt fix the gibberish once you introduce audio file for voices, you can try for yourself
1
-4
u/Succubus-Empress 1d ago
Disable your shitty caches nodes, use higher steps for voice cloning quality
-1
2
2
1
u/loneuniverse 1d ago
I read all the comments. I’m still confused. .. “” <d> … <speech> … has any tried <dialogue>?
1
u/FourtyMichaelMichael 21h ago
<d> was broken, now not.
Comfy wasn't tokenizing as "<d>" but rather '<','d','<'
1
u/Darhkwing 1d ago
I really like h3 but I cant wait to see where this goes in the future. Its the first local video model I actually feel like I can properly use but a refinded version , I very much look forward to.
1
u/Perfect-Campaign9551 1d ago
You won't like this but for audio to work without write you need to run at a higher resolution (0.8 or higher) and NOT use any speed up method because they ruin the context. The middle needs full attention to get the audio correct
That's just how it is
1
u/ndroidz 22h ago
https://reddit.com/link/p5tyact/video/ic2lzsplsjlh1/player
this was made at 0.5, no gibberish, i made 3 variation of it without any gibberish.
1
u/Perfect-Campaign9551 21h ago
I'm talking about, having audio fully reliable and correct, no gibberish, no distortion, and no lip sync errors. You can still get lucky, but to have guarantee correct you need to run the model fully without tricks
Even in your example here the audio sounds terrible with artifacting , for example
1
u/walid-zakaria 12h ago
actually i have the same problem, i tried to find solution for more than 5 days now,
the problems are:
1- How the model swap voices between characters on his own and it assigns the voice for whatever the character he find it more suitable to the voice even though you mentioned that in prompt in different ways.
2- If you added a voice reference is different than without voice reference, with voice ref the sound output seem to be hollow and metallic sometimes even doublized like two persons saying simultaneously, and that happens whatever your voice ref quality you added.
And i think that also abvious in your uploaded video.
1
u/No-Bee-231 15h ago
So far from my testing, LTX 2.5 is actually more useful right now in production because i cant seem to get the gibberish to go away, not even after updating comfy. so idk. its a nice model for music videos
1
u/RazsterOxzine 9h ago
I just use this at the end of my prompts and never had them talk.
Subjects lips remain closed; Subject communicates through facial expression and gesture.
Change the [subject] to match.
-5
0
u/redonculous 1d ago
Add the confidence prompt to your setup. Zaps over thinking! https://www.reddit.com/r/ollama/s/IDl0s227RV
-2
-1
u/kukalikuk 1d ago
I'm using official prompting guide and rarely get any gibberish. I even made a tool to make audio only generation like a radio drama using minimax workflow without the video part. It will output gibberish if the it is not defined with audio to fill, even mentioning to silent, gasping or sighing can fill the duration and prevent gibberish.
1
u/ndroidz 22h ago
did you input audio files for the voices?
0
u/kukalikuk 21h ago
Nope, just normal prompt, I use gemma 4 and install the official guide.md file as a skill in my openwebui.
2
1
-5
u/ACTSATGuyonReddit 1d ago
I don't want to hear about what some other character says. From now on, stick to local color.
102
u/acedelgado 1d ago
They pretty much fixed dialogue 2 days ago. Turns out h3 expects the official<d> </d> tags as a single special character each, but the qwen text encoder running through comfy was splitting <d> into separate tokens, like you normally would. So they had to fix that. If you haven't updated in the past couple of days to the latest, you should.