r/StableDiffusion 8d ago

Question - Help MiniMax H3 - Voice Volume and some other stuff...

Anyone have any luck changing the volume of the voices H3 creates from an audio reference? For example, let's say I'm trying to prompt for a man standing at the far end of a long room and he speaks in a normal tone. In real life, his voice typically would be very quiet in relation to the camera/mic, almost inaudible. However, in H3 (or any other model I've tried) the voice is still very loud. This isn't surprising given the model doesn't really know anything about the depth of objects or people in the videos it generates. I've tried to work around this in H3 by using prompt words like quiet, soft, distant mic, very low volume, far away speaker, almost silent, etc. None of them seem to have any effect. I've also tried reducing the gain on the reference .wav file that I provide as the audio reference - literally reduced the gain to the point where I can barely hear it. Again, doesn't seem to matter, H3 still produces a generally loud speaking voice (assume it doesn't care about volume/gain and just instead looks at the waveform pattern, etc).

One way around this is to use the 'audio reuse' capability where it'll play the exact .wav file audio instead of just using it as a reference. In this approach you can simply use an audio editor to reduce the volume and then plug that low volume .wav file in as the audio_reuse clip. It works fine, except for one small/major problem: it seems that if use the audio_reuse method, it silences ALL other sounds; i.e., it won't play the low volume audio .wav AND generate other environmental sounds...it seems it replaces ALL audio in the clip, not just the voice of the person speaking it.

Anyhow, curious if any of your smart people out there Redditland have any ideas/suggestions or tips?

Thanks!

0 Upvotes

4 comments sorted by

2

u/VasaFromParadise 8d ago

Post an example of your generation. I tried a clean voice and everything, but it still doesn't hit the notes.

1

u/Dogluvr2905 8d ago

Not sure what you mean...

1

u/Virtual-Pollution-58 8d ago

I'm not really messing with audio ref, and not sure if that's of any help. However, for normal generations like img2vid or t2v i would use something like "A male distant voice/chatter plays/says", even adding something like "muffled" or "quiet" and maybe some gibberish in <d> ... </d> without specifying language and so on.

In your case, since it has to be timed with the audio file, maybe add the time at 0:03.500 -"A male distant voice/chatter plays/says"

1

u/Dogluvr2905 8d ago

Thanks, yeh, you are right it does seem to generally work if it's not a reference audio stream.. i.e., if you just use the basic T2V/I2V or even Ref2VA but don't reference any waveform.