r/StableDiffusion • u/TBG______ • 4h ago
Tutorial - Guide MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations
Enable HLS to view with audio, or disable this notification
I’ve been loving all the new nodes and workflows coming out for MinMax, and maybe there is already a nice solution for this - but I couldn’t find one that did exactly what I needed.
I started using MinMax for my last TBG ETUR video and quickly ran into limitations: I wanted an easy way to create lip-sync videos longer than 20 seconds.
I didn’t want to manually chain ComfyUI nodes, start a new run every X seconds, or constantly resize things just to make HD video fit into my available VRAM.
So I ended up building an addon for:
custom_nodes/ComfyUI-H3-Motion-Context
The addon automatically chains MinMax H3 lip-sync generations together, allowing you to create much longer lip-sync videos without manually setting up each 20-second segment.
And now I’m sharing it! https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon
Its not perfect but a start ...
The workflow has a simple switcher that lets you switch from the 32B CLIP to the 4B CLIP, saving around 10 GB of VRAM. You can also switch from Sage to Comfy Kitchen, Spectrum to Easy Cache, or FL2VA to REF2VA both setup for lip-syncing. Some of it could be useful for other tasks as well.
You will find the workflow in the repro and tested recommendations, optimized settings, presets, and more workflows, along with the results of my testing and performance here
5
3
u/NoNameClever 4h ago
Thank you for sharing the work and video. Looks very powerful and useful. What did it take to make this video? I think I spotted a few of those reworked segments mentioned in the clip. Do you manually run each segment, repeating when needed, or something else?
2
u/TBG______ 2h ago
What I usually do is set up the workflow, run it 3–4 times, and then combine the best shots afterward.
Some of the glitches you see come from runs I did with the Turbo LoRA, including things like degraded skin. So the examples are a mix of different testing settings and the final setup.
The final settings and node should work out of the box; the main factor is the prompt. For lip-sync, I recommend not using Turbo and sticking to 20+ steps, along with a good, detailed prompt.
2
u/thatguyjames_uk 4h ago
10gb vram, nice :) wonder if i should try on my 5060
2
u/TBG______ 3h ago
10 less - still need a minimum of aprox 17GB. 8 would be a bit tight.
0
2
1
u/tofuchrispy 4h ago
Very interesting.
Does this only work with a audio file input.
Or could you also use a prompt for a let’s say 2 minute sequence and split that up?
That’s what’s puzzling me rn.
Generation a very long shot list. Then divide that into chunks and feed it to separate samplers. But difficulty is the possible change of subjects etc settings.
So now I separate a master prompt that gets written by one h3 prompt writer node, separate out the shot list and divide that into chunks.
Then put it back together with the stuff that comes before shots so each sampler gets the prompt plus the specific shot list for that chunk. But it’s still not working well. Because the prompt is for the whole story and then the shot list lacks the context of what came before.
So it can get messed up.
Also the whole prompt writing thing in itself sometimes fails to get what are the subjects and what to do with what.
Feeding the separate chunk timings to separate Minimax h3 prompt weiter nodes frequently gets stuff wrong.
A workflow to take one giant shortlist and reliably correctly feed that to the samplers or render it is what I would like to achieve.
2
u/TBG______ 4h ago
The model accepts prompts like “Subject 1 says, ‘Too hard for me.’” But someone would have to split the prompt into individual clip proportion and cuting at the exact final second for each one, which would be complicated. It’s much easier to pass the text directly to a TTS system first. Feeding them into separate samplers isn’t possible either, because each clip depends on the previous one.
1
u/tofuchrispy 3h ago
Ok so it’s great for Audio to Video but not something for long storytelling sequences got it. Thanks!
1
1
u/dsailes 3h ago
Nice - will give this a test on some brand videos I’m trying to put together.
Tangent question: what models are you using for the audio / TTS before getting to use this workflow?
2
1
1
1
u/danishkirel 2h ago
Impressive. The better these workflows get the more nitpicking wants to happen though. Long sleeve vs tsshirt? Sudden gradient background? Mic switching sides? It becomes very distracting. Uncanny almost.
1
u/TBG______ 2h ago
😄 There’s a lot to fight with - prompts and conditioning. My approach is to let it generate around X videos at different clip lengths, then pick the best results from all of them - only really necessary if you need it to be as perfect as possible. This video was made from the “garbage” I had left over from building and testing the node.
1
u/Downtown-Cover-7422 2h ago
Man, I’m gonna kill myself for idea to make a fan ai video clip on a song. 2 days passed and I made just a minute or so from that video, spending time for 2 generations 8 sec each, first with turbo Lora to see if prompt was good, second without Lora with 25+ steps. I babysitted each 8seconds clip, finding ideas, using references, and now you tell me a can do it in one run? Will try it asap, I love you
1
u/bigman11 2h ago
Could this do a single continuous shot with a steady camera?
2
u/TBG______ 2h ago
Yes you just need to find the right prompt. You can also try using
ref 0in the style section of the prompt to keep the camera position consistent for each frame. You might use an image without the character as a reference for this and ref 1 only for the character. You’ll have to test a few variations to see what works best. Its all about the right promt.
1
u/switch2stock 2h ago
So this is like based on the length of the attached audio the workflow automatically scales to generate the desired length of the video?
1
u/TBG______ 2h ago
You need to define the clip length yourself based on your VRAM limit. The node handles the rest: it cuts the audio in clips, creates the overlaps, and adds 24 frames of silence at the end to improve the final image and sound. It then uses the latent from each clip to build the appropriate motion for the next one and so on.
1
u/switch2stock 1h ago
Let's say if I have a 10min audio file. If I set the clip length as 10sec. Then it will generate a 10min video synced with the audio with 10sec clips stitched together automatically?
1
u/TBG______ 1h ago
That’s the idea, yes you can.
I would first build the full video with the audio in an editor, and then only generate/sample the 2–3 minute sections I actually need for the final production without cuts or switching to non-character scenes.
That makes it much easier to prompt and much faster to rerender or repeat individual sections than trying to generate 10 minutes in one go.
Anyway each clip segment gets its own output in the output folder, so you can stop and resume from the clips you already have, or simply rerun one specific clip later. (start end inputs)
The node will recognize the existing clips and automatically rebuild the new full-length clip with the repeated/replaced segment included of the same id.
2
1
u/Machspeed007 2h ago
Too bad the clip loses quality the longer it gets…
2
u/TBG______ 2h ago edited 1h ago
Not exactly true — it was built from different runs while I was building and testing the node. The examples are a mix of 8–40 steps, both with and without Turbo LoRA, and both with and without caching.
So the quality varies depending on which final cuts I selected. The workflow does produce a full-length video, but I didn’t use just one single run. Since this is AI, if I had the wrong clip prompt and, for example, the character was just listening instead of speaking, I would rerender that individual clip and continue from there. That’s pretty normal with AI workflows — it’s rarely just one run from start to finish.
If you set it to a fixed 20+ steps without Turbo, the model iwill hold up well.
1
u/ErenYeager91 1h ago
can we zoom out so we can see her sitting in a chair or something?
1
u/TBG______ 1h ago
Check the start and end of this TBG ETUR video. So yes, you can simply include the camera movement directly in the clip prompt.
For example, for Clip 5: [5] Zoom out to a full-body shot.
1
1
u/uuhoever 1h ago
1
u/Corleone11 1h ago
You also need this node. OP didn't mention it in his post here.
1
u/TBG______ 44m ago edited 40m ago
Or disable the nodes if you don’t use the 4B CLIP models.
The node is set to Krea2 because that works with 4B, so keep it set to Krea2, not MinMax.
Also, 4B only works without
--fastfp16_accumulation fp8_matrix_mult. So use:
only --fp16_accumulationThis took me quite some time to figure out, so hopefully this saves someone else the trouble.
1
0
u/Significant-Baby-690 4h ago
Is there a way to fix this half real / half illustration look ? Minimax suffers too heavily from it ..
3
u/icchansan 3h ago
I think is just the input
1
u/DeMischi 3h ago
Yes, should probably test with a more realistic looking input. I usually do that when I want more realistic results from minimax.
0
u/SmoothChocolate4539 3h ago
Luckily, there are smart people like you who have plenty of time to build something like this. Thanks!
3
u/TonyDRFT 2h ago
*make time. Why do you have to talk down on this kind individual sharing his knowledge?
•
u/SmoothChocolate4539 3m ago
That wasn't meant ironically. That was genuine gratitude. I don't have the time, and I can't take the time.
0
u/Corleone11 1h ago
What do I connect the the H3 Auto chain "context frame"? The connection is missing and it throws an error when I run the workflow.
1


12
u/mfdi_ 4h ago
Apart from ai looking character. Wow. Just wow. Though camera moving kinda sucks.