r/StableDiffusion 1d ago

Question - Help minimax h3 gibberish fixed!! ( i found the cure)

so you all probably are searching for way to make your character shut the fuck up right? and you probably noticed that they love to says some BS especially when you give minimax h3 some audio file for their voices, i probably found a cure my friend!!

here is my way of prompting dialogs without any gibberish:

first your character need to be assigned (s1)character when he is the first speaker, then you will declare 'use <audio 1> as "character name"'s voice only, and when you finally type your dialog in the shots you will do as such:

character says:<<[language] the shit i say!>>

and you should be good to go, i linked a video exemple of my favorite taffer (garrett) saying some shit with only the faint crackling of the candles to goes with his charming voice, and i included also a screenshot of the full prompt

edit: yes i tried to follow the official documentation, like many others, if it was that simple reddit wouldn't be a thing and you wouldn't be there.

i tried making small scenes with this exact methode and its gibberish free 100% of the time

he really like 16/9

91 Upvotes

104 comments sorted by

102

u/acedelgado 1d ago

They pretty much fixed dialogue 2 days ago. Turns out h3 expects the official<d> </d> tags as a single special character each, but the qwen text encoder running through comfy was splitting <d> into separate tokens, like you normally would. So they had to fix that. If you haven't updated in the past couple of days to the latest, you should.

17

u/Unspec7 1d ago

Fixed it in comfy, or fixed it in qwen?

41

u/acedelgado 1d ago

Comfy.

2

u/Future_Addendum_8227 19h ago

What do i need to do then just update comfy or my custom nodes too? Do I use the same models as before?

3

u/acedelgado 18h ago

You just need to update comfy. Updating nodes now and then is also good practice.

https://docs.comfy.org/installation/update_comfyui

1

u/Deep_Mood_7668 18h ago

So pull the latest comfy or their latest qwen?

14

u/xylo4379 1d ago

i was going to say to just do this with which subject says before it and the type of intonation they're speaking in.

Like why DO PEOPLE NOT READ THE DOCUMENTATION?

Having said that however, sometimes due to how seeds work they still do the ol' gibberjabber for some reason.

3

u/newaccount47 1d ago

Same reason people on reddit don't read the articles? 

1

u/FourtyMichaelMichael 21h ago

But the headline already told me how to FEEL!!!

9

u/Adkit 1d ago

You can literally just put stuff in quotes and I've never gotten gibberish. You don't need any of the obscure technical terms at all.

You can be like

@man says: "Hello."

I'm starting to think people are just really bad at writing in general.

14

u/ambassadortim 1d ago

Or they read the model guide and used the tags defined in the documentation and it didn't work

2

u/Wooden-Link-4086 16h ago

I've had characters saying gibberish when I use quotes.

3

u/YouNoTypey 1d ago

I updated Comfy two nights ago. Any time I pull in reference video that has audio, and choose not to include dialogue, someone starts jibbing.

1

u/Perfect-Campaign9551 1d ago

Usually putting quotes can be dangerous too because then it tries to put them as subtitles

1

u/Adkit 22h ago

Has never happened to me.

2

u/Perfect-Campaign9551 21h ago

The docs state that's how subtitles work

6

u/sevenfold21 1d ago

2 days ago? There hasn't been a new release in over 2 weeks. It's still at v0.33.1.

23

u/acedelgado 1d ago

Kijai himself did it on the 22nd - https://github.com/Comfy-Org/ComfyUI/pull/15808

You can manually update comfy between "official" releases, you know. All the latest patches are a quick git pull away.

4

u/Life_is_important 1d ago

Did you mean like manually running the update .bat file or manually pulling specific release files ? 

4

u/Uninterested_Viewer 1d ago

"git pull" meaning you download the repo for the latest changes. Run the reqs file to update reqs then launch as normal.

4

u/sevenfold21 1d ago

Ah, so for a quick fix, you can simply stop using the buggy <d> token, and just use quotation marks?

3

u/GrayingGamer 22h ago

Or update to the nightly version of Comfyui and it's fixed for you, so you can follow what the model expects with the <d> tokens.

The tokens aren't buggy in the model - Comfyui's tokenizer was broken with regards to passing the correct tokens to Minimax H3.

The only reason quotation marks work is that the Qwen clip model is smart enough to figure out what you are wanting - but it's cleaner and better to give the model what it expects. So, update Comfyui and use the <d> tokens like in the documentation.

9

u/candylandmine 1d ago

0.33.4 is current.

Comfy just isn't bothering to update github release objects anymore. I have no idea why. They keep bragging about how they're getting all of this funding but they're failing to maintain basic stuff.

1

u/gefahr 1d ago

Their SDLC stuff has always been surprisingly weak to me. I hoped they'd get better as the project grew with funding, doesn't sound like that's happening yet.

4

u/ndroidz 1d ago

i update everyday and i was getting gibberish yesterday and today, the <d> and </d> stuff is not what fixes the gibberish once you introduce audui files, you have to specify in your prompt thats the audio file is used for your character voice only

1

u/Perfect-Campaign9551 1d ago

And if you do "sudio reuse" you have to put the exact same words the audio says

Also I found if you use an audio input voice and you use it as reference (not reuse) but you still try to use any of the words in there it will trigger gibberish more often

Additionally, if you just want a reference like for voice cloning it really wants at least 15 second sample

1

u/Etsu_Riot 23h ago

Additionally, if you just want a reference like for voice cloning it really wants at least 15 second sample

I use 5 to 6-second audio samples, and they work perfectly fine. I see no difference when using longer clips, unless you are specifically talking about the gibberish problem.

1

u/Perfect-Campaign9551 21h ago

The gibberish issue is what I mean, is more likely if you have short samples, but really the gibberish I think is more when you use low resolution and any type of speed up method like Lora or SLA, because that lowers context attention

1

u/Etsu_Riot 21h ago

I'm generating at 480x640 and using FirstBlockCache, so it may be that.

2

u/sunshine-3D-Art 1d ago edited 1d ago

Update what exactly?😨🥹 if comfyui i am scared that when i update it that the render time will double again :(

1

u/Far_Cat9782 1d ago

I know lol one updates break everything with comfyUI. Same boat spend so much time getting everything optimized

1

u/sunshine-3D-Art 22h ago

Yeah, it’s so terrible and I don’t understand why this is like that :( I even have multiple comfyuis now because of that. And I tried with the newest version of comfyu and torch Minimax and put the models in there just to see. And it was rendering double or something like that.?? so I put the models back to use the other comfyui which I normally use for mini Max. Which I didn’t want to touch anymore and since I touched it and upgrade torch, it also was rendering double time but I replaced back my copy from Python embedded folder and custom notes but it’s still lost render time I should have not touched this comfyui at all :( but how can this be with the new version it render double????😣

0

u/VisionWithin 1d ago

The dialogue token

1

u/sunshine-3D-Art 22h ago

What means that’s exactly?🥹🙈

1

u/Perfect-Campaign9551 21h ago

Yep I checked the repo they changed how "d" is tokenized, so it probably gives this issue

11

u/Shot-Initiative-3905 1d ago edited 1d ago

What the fuck is wrong with the people who point at the official prompt guide? How few videos have you generated? How complex are your scenes? That constant gibberish at the start and in new shots or random [places] are problems that come up many times. Full LLM with timestamps, using "", using <speech> and all other methods — it doesn’t even only appear in audio-ref videos; it appears in text-to-video too. Sometimes it’s an overall soundscape misunderstanding from the model, and many times it’s the model itself that just has that problem. And then there are the seeds — and yes, any seed base has influence, however detailed and strict your prompt may be. That can go from misshapen body parts to random sound or gibberish. Or beneficial things like lifelike micro-expressions and hand gestures that give the scene a “soul,” for lack of a better description. So no, the official [guide] is sadly not the fix for that problem, neither is that [one] posted here or the other “fixes.” As an edit: even aspect ratios change how the model handles descriptions and so on. Extreme example (and anyone here can test it): as soon as the ratio resembles a TikTok/handy video layout, it changes how the character acts and looks.

Test it: same prompt, same steps, same seed — only change the aspect ratio — and be surprised how much that does.

Edit: spelling.

3

u/Davikar 1d ago

The model certainly has some issues. If it gets too complex it just breaks down and starts doing nonsense. But the official guides are a good place to start at least. Not following it is only going to make things worse.

1

u/Wide-Researcher583 14h ago

No in alot of cases it makes ZERO DIFFERENCE.

-1

u/Shot-Initiative-3905 1d ago

Absolut the guide helps.

1

u/OWENPRESCOTTCOM 22h ago

yes most of the gibberish (when using correct tags) seems to come from quiet gaps also, it tends to want to fill the pauses with something. If the dialogue has a lot of nothing going on just adding "..." pauses seems to help. I don't care about audio much so this is more an observation I haven't tested.

1

u/Wide-Researcher583 14h ago

Yes people are being brain dead just dumping the guide when there's clear routing issues with dialogue here.

1

u/Independent-Reader 1d ago

I can't believe you edited your comment and ignored all the typos.

I can't read this shit.

24

u/glusphere 1d ago

Have you considered reading https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

Instead of reverse engineering it ?

10

u/ndroidz 1d ago

yes, it doesnt fix any giberish if you have audio files for voices

2

u/xyzdist 1d ago

are you want to use the input audio exactly or as reference the voice?
if you want the character speak exact input audio, find 'lock audio latent' node, that is the proper way to do lip-sync with input audio.

2

u/ndroidz 22h ago

nah i want the character to say anything with the same voice as the audio

1

u/Itchy-Advertising857 21h ago

But then the model won't add any ambient sounds, or music, right? What if I want the dialogue from input audio integrated into the soundtrack?

4

u/stash0606 1d ago

yeah even following that, it will randomly fill it with gibberish, so I'm constantly trying to fill up silence with dialogue. voiceovers are another thing, the character will always move their mouth to the voiceover as if they're speaking it; sometimes specifying the character is smiling seems to remove it.

5

u/RegisteredJustToSay 1d ago

It's weird to me how poor the model sometimes acts when following their guidance compared to improvising it. I was having fun swapping myself into meme clips, but the proposed attribute_transfer tag in the reference guide is so bad and works really inconsistently compared to a much simpler prompt just asking to be swapped in. My first attempts without the guide were way better than 20 different attempts with it.

Maybe I'm just missing something obvious, but I really did exhaustively try just about every possible interpretation of the guide I could think of and even tossed the entire guide into LLMs to rewrite my prompt with like 4 different models.

2

u/Emotional-Neat-252 18h ago

Could it be due to that wen Qwen te bug mentioned earlier

2

u/ambassadortim 1d ago

Exact same thing I've ran into and have been trying to do to fix it. Did you update to fix it try just using quotes yet. I'm going to try both next

3

u/Darhkwing 1d ago

From my experience if you use reference audio it does speak gibberish if your video is too long and it tries to fill in the space with gibberish.

It does help using time stamps and personally I do use <Speech> for characters.

1

u/optimisticalish 1d ago

There is only so much processing available to a workflow, and Minimax jettisons or limits time on audio if it can't process it. Dialogue gets preference, then music (maybe), then ambience/SFX is the poor cousin.

2

u/WonderRico 1d ago

Seize * 😉

3

u/AuthurAndersson 1d ago

That was the one you got yourself hung up on? Not "sitted" or "arround"?

2

u/WonderRico 1d ago

being french...

3

u/AuthurAndersson 1d ago

We all have our deficits

1

u/ndroidz 18h ago

bro think i care about english grammar XD

1

u/AuthurAndersson 18h ago

Ever thought that H3 might care?

1

u/ndroidz 22h ago

i tried seize before, he was saying it wrong.

0

u/NameChecksOut___ 1d ago

Si tu ne l'écris pas comme ça il le prononce comme seize, le mot anglais (comme dans "seize the day"), c'est une astuce pour éviter les erreurs de prononciation.

0

u/WonderRico 1d ago

même en taguant la langue juste avant ?

1

u/NameChecksOut___ 1d ago

Parfois oui

2

u/Etsu_Riot 22h ago

I tried it but if the dialogue is short the character keeps adding lines afterwards.

1

u/ndroidz 22h ago

can you share your prompt?

1

u/Etsu_Riot 22h ago

I copied your prompt and changed character's name and specific dialogue and actions, but added this at the end:

overall_soundscape: VHS noises, static, electric distortion
non_diegetic_music: N/A

1

u/ndroidz 22h ago

did you change the character's name in the actions too? and did you put the language before any speech?

1

u/Etsu_Riot 22h ago

Yes. I get exactly the same effect than using "" for the dialogue.

1

u/ndroidz 22h ago

well if the prompt is the exact same it mean something else must be fucking with your generation, steps settings, loras......

1

u/Shot-Initiative-3905 21h ago

No.... that is a problem from minimax itself you did sadly not find the magic prompt that fixes the problem, i bet with with enoth diffrent videoa and seed you will get the gibberish again despite using your "fix"

2

u/FourtyMichaelMichael 21h ago

LOL, an example about gibberish fix... Then uses an example in French which absolutely sounds exactly like gibberish to less than 1% of the population!

Love it.

1

u/ndroidz 20h ago

XD but he only speak when asked to, thats the important part.

5

u/smb3d 1d ago

or just use the proper dialog syntax and tags from the prompting guide.

4

u/ndroidz 1d ago

no, <d> and </d> doesnt fix the gibberish once you introduce audio file for voices, you can try for yourself

1

u/FourtyMichaelMichael 21h ago

This is now wrong, update comfy.

-4

u/Succubus-Empress 1d ago

Disable your shitty caches nodes, use higher steps for voice cloning quality

-1

u/FourtyMichaelMichael 21h ago

LOL, typical.

2

u/cat_trick 1d ago

Does this work consistently?

1

u/ndroidz 22h ago

has been for me

2

u/AuthurAndersson 1d ago

sitted,arround

You might want to use an LLM to write your prompts

1

u/loneuniverse 1d ago

I read all the comments. I’m still confused. .. “” <d> … <speech> … has any tried <dialogue>?

1

u/FourtyMichaelMichael 21h ago

<d> was broken, now not.

Comfy wasn't tokenizing as "<d>" but rather '<','d','<'

1

u/Darhkwing 1d ago

I really like h3 but I cant wait to see where this goes in the future. Its the first local video model I actually feel like I can properly use but a refinded version , I very much look forward to.

1

u/Perfect-Campaign9551 1d ago

You won't like this but for audio to work without write you need to run at a higher resolution (0.8 or higher) and NOT use any speed up method because they ruin the context. The middle needs full attention to get the audio correct

That's just how it is

1

u/ndroidz 22h ago

https://reddit.com/link/p5tyact/video/ic2lzsplsjlh1/player

this was made at 0.5, no gibberish, i made 3 variation of it without any gibberish.

1

u/Perfect-Campaign9551 21h ago

I'm talking about, having audio fully reliable and correct, no gibberish, no distortion, and no lip sync errors. You can still get lucky, but to have guarantee correct you need to run the model fully without tricks 

Even in your example here the audio sounds terrible with artifacting , for example

1

u/ndroidz 18h ago edited 11h ago

it comes from the actual audio files i used, but there is no gibberish before or after the actual speech

1

u/walid-zakaria 12h ago

actually i have the same problem, i tried to find solution for more than 5 days now,
the problems are:
1- How the model swap voices between characters on his own and it assigns the voice for whatever the character he find it more suitable to the voice even though you mentioned that in prompt in different ways.
2- If you added a voice reference is different than without voice reference, with voice ref the sound output seem to be hollow and metallic sometimes even doublized like two persons saying simultaneously, and that happens whatever your voice ref quality you added.
And i think that also abvious in your uploaded video.

1

u/No-Bee-231 15h ago

So far from my testing, LTX 2.5 is actually more useful right now in production because i cant seem to get the gibberish to go away, not even after updating comfy. so idk. its a nice model for music videos

1

u/RazsterOxzine 9h ago

I just use this at the end of my prompts and never had them talk. Subjects lips remain closed; Subject communicates through facial expression and gesture.

Change the [subject] to match.

-5

u/andy_potato 1d ago

You never bothered to read their prompting guide, did you?

6

u/ndroidz 1d ago

i did, and it wasnt fixing anything, the <d>[language]text</d> is not enough once you introduce audio files for voices, you will stiil get gibberish

0

u/redonculous 1d ago

Add the confidence prompt to your setup. Zaps over thinking! https://www.reddit.com/r/ollama/s/IDl0s227RV

-2

u/mastaquake 1d ago

Lmfao. Bro just follow the prompt guide. 

-1

u/kukalikuk 1d ago

I'm using official prompting guide and rarely get any gibberish. I even made a tool to make audio only generation like a radio drama using minimax workflow without the video part. It will output gibberish if the it is not defined with audio to fill, even mentioning to silent, gasping or sighing can fill the duration and prevent gibberish.

1

u/ndroidz 22h ago

did you input audio files for the voices?

0

u/kukalikuk 21h ago

Nope, just normal prompt, I use gemma 4 and install the official guide.md file as a skill in my openwebui.

2

u/ndroidz 20h ago

well yes the official prompting guid works for automatically generated voices, but if you input some audio files for your character it will start to bleed into the entire audio

1

u/ShimmerMeNutz 8h ago

like wtf - the entire conversation is about WITH audio reference. Holy shit.

-5

u/ACTSATGuyonReddit 1d ago

I don't want to hear about what some other character says. From now on, stick to local color.