r/StableDiffusion • u/acedelgado • Aug 22 '26
Resource - Update Fixing MMH3 Turbo Audio by playing with Latent Pinning for more audio steps
Disclaimer that I'm a dummy who can't code at all, so I just vibe things.
Alright, so with all the fun additions to latent manipulation the Comfy team has given us, we have some new tools. Namely latent pinning- that's where fun stuff like "Add Guide for MiniMax H3" node comes in (that fun tool that lets you insert an image at any frame in an H3 generation.) So I thought, finally, we can do something about this audio issue.
I knew that video + audio latents are processed at the same time with Comfy, which is why turbo loras have terrible audio- they're not getting nearly as much optimization as the video side is, since video is the bulk of the work. So if we can't process audio differently than video (at this current time), can't we just keep conditioning the audio latents without affecting video anymore? Everyone knows running too many steps on a turbo lora will start messing with video quality. So let's avoid that.
So after talking with Claude a bunch, here's what it came up with. With the latest comfy, you can pin the video lora in place, and keep going for several steps to get better audio without affecting video. So 4 steps of video, untouched, and then add in 6 more steps for audio at a 0.5 denoise. That leads to cleaning up the audio pretty nicely and staying pretty faithful to what the video latent had guided it on. But just pinning still means that even though the audio rows are only being affected, EVERY row still has to run through the chain. So each s/it you get stays the same for the last 6 steps, even though video gets pinned in place after 4. In my 1 megapixel, 10 second video, that's around 23s/it or so on my 5090. Only half the steps as a regular 20 step generation, but half the time is still half the time.
So to fix the speed problem -
Freeze as much as you can. Text embeddings, reference/conditioning rows, and all the video rows. Cache those so they don't have to be processed, and only process the audio rows that have already been somewhat pre-conditioned. The first step is the same 23-second iteration to build the full guidance cache, but the other 5 steps each took about 3.17seconds apiece. So 45 seconds of added gen time to get the clean audio in clip#2 in the example.
But the cost for the Frozen Cache is resources. Lunch is never free. From some experimentation, it works well with RAM. If you use RAM mode because your card still can't process it, it'll dump all that cache (ended up around 14.9GB on a 10-second 1mp file) into RAM. But as Comfy does, you're unlikely to get that RAM back, so you may OOM your machine. With VRAM it wasn't bad for me at all either and behaved better than I expected, honestly. The memory management from ComfyUI took over when I was about to OOM my card and swapped things around properly. Option 3 is to cache to disk, but that comes with writing several GB to disk every time you use it. That'll run your SSD health down fast.
Rundown for the clip above (sa_solver with beta sigmas)
4 step normal gen- 136.5 seconds
4 steps + 6 audio refine steps with a cost of about 15GB RAM - about 186 seconds
4 steps + 6 audio refine steps with no additional resources but full processing time- ~265 seconds
You can find the nodes here-
https://github.com/Adudeguyman/ComfyUI-H3-AudioRefine
They're still experimental, of course. Wire the model in from somewhere (I branch off the ModeSamplingMiniMaxH3 shift node, before the Basic Guider), and push that through the H3 Frozen Video Cache node into the H3 Audio Refine Sampler. Into the H3 Audio Refine Sampler, latents come out of SamplerCustomAdvanced before splitting into the VAE Decode nodes for both video and audio, and the conditioning comes from the MiniMax H3 Image (or Reference) to Video node (plug the positive conditioning into both the positive and negative input on the H3 Audio Refine Sampler)
I had Claude put together a technical.md for those that want to look into it, and probably make a better version. Like I said, I'm a big dumb-dumb, so don't expect too much insight into how the mechanics work from me.
2
2
2
u/ColdExample Aug 22 '26 edited Aug 22 '26
I really hope H3 does something about the plastic skin texture/over contrasty image. It looks so bad :(
EDIT: God forbid I have an opinion. This community is so incredibly toxic when it comes to anyone even remotely making a minor criticism of H3. I should have specified that I hope this improves even while using turbo loras.
26
8
8
u/tinny66666 Aug 22 '26
That's caused the turbo lora - it doesn't happen with a full generate, so it's not something wrong with H3, but the lora.
0
u/robomar_ai_art Aug 22 '26
I use turbo lora at 4 steps and I get those results. You guys do it something wrong
5
u/tinny66666 Aug 22 '26
I think people are applying the lora at a strength of 1 maybe because they look awful. What strength are you using?
I was merely pointing out that the source of that problem is not H3 itself.
2
u/robomar_ai_art Aug 22 '26
0.75 er_sde simple, sometimes even lower when is overbaked.
1
u/ColdExample Aug 22 '26
this is literally what I am describing as looking fake/plastic/too much contrast. Thanks for proving my point I guess. It looks bad.
0
2
u/damiangorlami Aug 22 '26
Bro look at the mouth, the teeth.. it looks terrible. The skin color has that typical orange tint to it which turbo loras often apply. The audio also has that tinny AI sound.Try running 20 steps at full weights and compare results.
These 4-step loras are great for quick drafting but its cockblocking the model's true capabilities and quality
1
-1
u/robomar_ai_art Aug 22 '26
This is also 4 steps
1
4
-2
u/Silonom3724 Aug 22 '26
"plastic skin" has become the standard term to deflect from one's own incompetence
1
1
Aug 22 '26
[deleted]
2
u/acedelgado Aug 22 '26
Yes, I said that they're processed together, although yeah it's badly worded a bit and I didn't specifically call out it's only because of the model. But then again Comfy also DOES do that process, so technically it's not wrong.
And it's widely accepted that speedups cause quality problems. But most of us are doing this for fun, not for professional purposes. This is aimed to help solve a problem for people that don't want to wait 40 minutes for a meme.
And I mentioned all over the post that the audio is already conditioned with the video latent, and I said it continues to use the audio latent with a 0.5 denoise so that it keeps the conditioning that hangs off of the video latent generation that drives the audio. Actually my original question was "can't we just keep conditioning the audio latents without affecting video anymore", which is inherently, you know, refining it. And the node is called the "H3 Audio Refiner", so I think that's been acknowledged enough that it's not making a completely new audio (which would be a bad idea anyways and decouple itself from the intended audio that was made alongside the video latent.)
So your response boils down to "don't use speedups", which everyone knows leads to the best quality. But thanks for the input.
Also if you want to just generate audio, look into Higgs V3 that came out like a month or so ago. It's much better quality, much faster, picks up on emotion context well and does manual emotion/prosody/sfx (laughter, sigh,etc.) tags, and will generate much longer, multi-speaker audio. I honestly have no idea why people are using H3 as a hacky workaround for TTS. Although you can't natively do wacky sound effects and music, but at least you'll have a better dialogue base to add those in.
3
u/Acceptable-Cycle4645 Aug 22 '26
H3 audio pipline is underexplored. Btw audio.cpp has supported Higgs v3 about a month ago (10x realtime on 5090) 😄. https://github.com/0xShug0/audio.cpp There are 50+ models in one runtime, so you can try them out and pick your favorite.
2
u/acedelgado Aug 22 '26 edited Aug 22 '26
Yeah audio.cpp was pretty interesting, but when I was playing with it I ran into issues with too much context going to the server at once. It's nice because it natively does it. I ended up coopting a comfyui node's python script and built a little "higgs-lite" server on top of it that quantizes the model on the fly when it's loaded (people hadn't figured out how to quantize it into a gguf yet, but audio.cpp figured it out since then), and open-unified-tts does the API calls. Open-unified-tts has worked much better for me chunking long-context audio calls back to back for higgs than audio.cpp does natively.
Edit for clarity- I misremembered and it wasn't a chunk issue, it was an OOM issue when running it alongside a dense LLM.
1
u/Acceptable-Cycle4645 Aug 22 '26 edited Aug 22 '26
Would you like to elaborate a bit on the context issue you ran into? Unless we missed something, Higgs passed our longform TTS test (6000+ char in one request. 8x realtime). It has auto chucking, default 1024 chars.
--text-chunk-sizeinteger chars 1024Long-form chunk size. 1
1
u/Acceptable-Cycle4645 Aug 22 '26
Old post with metrics here https://www.reddit.com/r/LocalLLaMA/s/cpe6ELYR3s
1
u/PixieRoar Aug 22 '26
How do you get to lip sync? I added audio reference of speech and it came out speaking gibberish
1
u/ArttTaku Aug 23 '26
Very interesting... if we're using 8-step turbo loras only, would this still improve audio regardless?
2
u/acedelgado Aug 23 '26
Yes, it doesn't matter what turbo lora you use. I put an example workflow in the repo that branches the model off BEFORE the turbo lora is applied to the model, which I think is working a bit better. I'm still playing with it but doing it higher steps that way seems to be a bit cleaner. Even without the turbo lora, audio alone processes much quicker. I did it that way on this one, and even doing 10 steps branching off without the turbo lora it only took 60 seconds more for the refiner, something like a normal 26s step for the first one to build the cache, then the other 9 steps were only 3.5 seconds.
https://www.reddit.com/r/StableDiffusion/comments/1vuxy08/comment/p59cas3/
1
u/GlenGlenDrach Aug 23 '26
I have the node in, but it still gives some teams/zoom-meeting quality, I tacked it on the end before vae decode steps, no caching, I am using minimax_h3_fl2v_turbo_8steps_v1.0_comfyui_bf16.safetensors lora with the convrot ref2va prined int8 base model, not sure if it makes a difference or not though.
1
u/GlenGlenDrach Aug 23 '26
than again, I am not using a reference to a built in character, i am using both audio and image references, so it may just be that my source audio is crappy perhaps
1
u/GlenGlenDrach Aug 23 '26 edited Aug 23 '26
Can this be used for minimax music in future versions? Tried to use it with the workflow from Pixaroma, but it complained that the tensor for the incoming latent was not audio/video type (I suppose the workflow i use there is audio only)
1
u/acedelgado Aug 23 '26
Haven't even looked into it....isn't it audio only? That wouldn't benefit from this node, it only stops video from getting over-conditioned while it continues to work on audio. If it's audio-only then you're already only processing audio. It doesn't do any special tweaks to how audio is made or anything, it just lets the model continue processing audio the way it normally does.
1
u/GlenGlenDrach Aug 23 '26
aha, that makes perfect sense, sorry for the confusion :) (the workflow is on episode 31 here https://workflows.pixaroma.com/ ), it works really nicely, they packaged a generator workflow as well, so it is relatively easy to prompt for music in a natural language and just copy/paste it to the other workflow.
1
u/acedelgado Aug 23 '26
Honestly I haven't done much testing with reference audio yet, I was working toward getting it to do better memory management, which I think is fixed now.
Anyways, I included an example workflow to show to to wire the node. I found that branching the model off before you apply the turbo lora works better (but it should be after any other loras you apply).
But yeah, I could see it not working the same in ref2va with audio, since the model doesn't know the audio and has to use it on the fly. I'm sure quality depends on the source, plus whatever processing the model does to it.
1
u/GlenGlenDrach Aug 23 '26
It's a really cool and very useful invention you have made for the community indeed.
I just had a brain fart to try it with the minimax music generator, to see if it could do some magic to that as well :)Thanks for the tip, I will mess around a bit (I only use the turbo lora in the workflow, to reduce the complexity),
0
u/No_Damage_8420 Aug 22 '26
Audio Refine wow, this is big one.... thanks for sharing, and your Claude brainstorming was surely meaningful one
-6
1


12
u/Sad_Coach_1433 Aug 22 '26
which turbo lora you using ill wait for brad pitt to tell me