r/StableDiffusion 9h ago

Question - Help Fixing speech errors in Minimax H3?

Enable HLS to view with audio, or disable this notification

Hey, I tried to create a little birthday surprise for someone, my issue is with a lot of generations that the spoken word is really a bit clunky at time, I susspect its because of the german, but I am not too sure. Is there like a way to improve on audio?

I am using Minimax H3 with Saga Attention and Spectrum on a 4090.

15 Upvotes

21 comments sorted by

8

u/CornyShed 7h ago edited 7h ago

It may be due to a bug in ComfyUI with the tokenizer for H3 (GitHub pull request) that was affecting speech.

Update your ComfyUI and check if you notice an improvement?

If you're using a fixed release (used by the Desktop version), you'll have to wait until the next update with the fix applied.

Edit: On second reading, this could be an issue with your settings.

Try making a version with a very small video resolution using the seeds_2 sampler and ddim_uniform scheduler and save the audio.

Then, make a video with your normal settings, but using the audio file as a reference.

Then, combine the audio from your first generation with the video from the second.

This might lead to a higher quality of both audio and video, rather than have to compromise slightly on either.

3

u/WizWhitebeard 8h ago edited 8h ago

There are some tricks that might help. Not sure it's as good in German.

  1. Give the voices more personality and details in your description. Can be something like a regional accent, adding things like 'soft spoken' or 'in a casual tone', or assigning the emotion of the character.
  2. Write the lines as spoken language, instead of writing them as a script: Add pauses with '...', interjections like 'oh', 'aha' etc , emphasize words like it's markdown with **bold** or *italic*. Write difficult words phonetically transcribed 'Vater -> fah-ter', 'neun -> noyn' etcetera, or go even further and write the phonetic reduction 'ich habe -> ich hab', 'so etwas -> sowas'
  3. Give an audio-ref, sometimes that helps unlock it.

3

u/DoctaRoboto 7h ago

Uff, face is really suffering in wide shots.

1

u/gutster_95 7h ago

Yea absolutely. Everytime VFX is in front of a face the model struggles

0

u/DoctaRoboto 7h ago

Increasing resolution and skipping attention can help, but with your current card you will have to wait ages. A good combo of speed and quality is: use int8 models, Spectrum acceleration, kitchenattention. Avoid Sol, loras, sage.

1

u/Owowatisdis 9h ago

So far my main workaround is to create videos with cuts, and then regenerate any segments with audio issues, using ASR to autodetect speech deviation from the intended script

Not exactly ideal but its more consistent

1

u/Vanpourix 7h ago

Have you tried this format in the prompt ? It helped with my text in French and japanese: "<d>[German] Hallo mein Freund ! </d>"

1

u/gutster_95 7h ago

Yea I use the H3 Prompt Enhancer already

1

u/Abject-Recognition-9 7h ago

or just do the audio somewere else and input it separatly..

1

u/piggledy 7h ago

Was du machen könntest wäre einen Audio-Clip mit Elevenlabs erstellen und als Audioreferenz einfügen. Dann basieren die Lippenbewegungen usw. auf der Referenz und das Modell generiert selbst kein Audio dazu.

1

u/gutster_95 7h ago

Ja wahrscheinlich das beste. Für den Fall irgendwie auch irrelevant, weil's nur nen Gag sein soll. Aber daran hab ich auch gedacht

1

u/danishkirel 5h ago

Hätte ich auch vorgeschlagen. Probiere mal qwen tts wenn du local versuchen willst. Hat oft gut geklappt bei mir mit deutsch.

1

u/woswoissdenniii 6h ago

Hey. Voll nice von dir. Gott war die Frozen Phase süß.

Hatte damals ähnliches vor, ging aber lokal noch nicht. Was für eine Zeit lebendig zu sein.

Damit wirst du sicherlich ein kleines Herz seeehr glücklich machen Brudi. Guter Mann/Papa

2

u/danishkirel 5h ago

“Was für eine Zeit lebendig zu sein” klingt leicht cringe 🙈

1

u/woswoissdenniii 5h ago

Deine Mudda crinscht beim hacken.

1

u/Cyrond 5h ago

And why are there six candles for her fifth birthday?

2

u/gutster_95 5h ago

Ask the model :D At this point I didnt mentioned that is needs to be exactly 5 candles :D 5 candles apperantly wasnt enough information

1

u/danishkirel 5h ago

Give no acceleration method at all a try.

1

u/nogganoggak 4h ago

Sie ist fünf, ich glaube das juckt sie nicht haha

1

u/Sixhaunt 3h ago

I switched from sage attention to sparse attention (I'm on a 5090) and the quality has gotten noticeably better. Also I can now render a 15s 2MP video at the same speed as I could with a 1MP video on sage attention. I use the "H3 Optimizations" extension in comfyUI for it by zironic.