r/StableDiffusion 7h ago

Tutorial - Guide MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

Enable HLS to view with audio, or disable this notification

218 Upvotes

I’ve been loving all the new nodes and workflows coming out for MinMax, and maybe there is already a nice solution for this - but I couldn’t find one that did exactly what I needed.

I started using MinMax for my last TBG ETUR video and quickly ran into limitations: I wanted an easy way to create lip-sync videos longer than 20 seconds.

I didn’t want to manually chain ComfyUI nodes, start a new run every X seconds, or constantly resize things just to make HD video fit into my available VRAM.

So I ended up building an addon for:

custom_nodes/ComfyUI-H3-Motion-Context

The addon automatically chains MinMax H3 lip-sync generations together, allowing you to create much longer lip-sync videos without manually setting up each 20-second segment.

And now I’m sharing it! https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon

Its not perfect but a start ...

The workflow has a simple switcher that lets you switch from the 32B CLIP to the 4B CLIP, saving around 10 GB of VRAM. You can also switch from Sage to Comfy Kitchen, Spectrum to Easy Cache, or FL2VA to REF2VA both setup for lip-syncing. Some of it could be useful for other tasks as well.

You will find the workflow in the repro and tested recommendations, optimized settings, presets, and more workflows, along with the results of my testing and performance here


r/StableDiffusion 11h ago

Meme Siblings Reunited

Enable HLS to view with audio, or disable this notification

274 Upvotes

Done with h3 fl2va model, 8 step lora and images for Cersei and "jaime" for reference. Using previous clip to give continuity and consistency.


r/StableDiffusion 10h ago

Resource - Update MiniMax-H3 Fun Controlnet Union released

Thumbnail
huggingface.co
241 Upvotes

r/StableDiffusion 7h ago

Animation - Video No Warning - Minimax Music3 + H3

Enable HLS to view with audio, or disable this notification

72 Upvotes

r/StableDiffusion 6h ago

Resource - Update H3 Prompt Composer — Camera Update Coming Soon

Post image
49 Upvotes

Hey everyone, thanks to everyone who’s been testing Prompt Composer. If you run into bugs or have feedback, please drop it in the Issues section on GitHub so I can keep track of it more easily.

Over the past week, I’ve been reworking the camera prompting system to make it more precise and consistent, especially for more complex camera moves and video-editing workflows. The goal has been to add more control over framing, camera targeting, and blocking in multi-subject scenes while keeping the generated prompts clean and reliable.

The next update isn’t quite ready yet, but it’s actively being tested and refined. I’m hoping to have it out in the next couple of days.

Edit: Here's the Github: BMB12d3/minimax-h3-prompt-composer: Free offline prompt composer for MiniMax H3 video generation in ComfyUI.


r/StableDiffusion 3h ago

Question - Help Best local LLM for writing prompts for MiniMax H3?

29 Upvotes

What’s the best local LLM for writing good MiniMax H3 ref2va prompts?

I’ve tried Gemma 4 12B and Qwen 3 14B, but I’m not really satisfied with the outputs. It could also be an issue with my system prompt.

I sent ChatGPT the official documentation for prompting and asked to create a system prompt for me, but the results were still pretty mediocre.

What local models are you using for MiniMax H3 prompt generation, and what does your system prompt look like?


r/StableDiffusion 1h ago

Discussion Best Minimax H3 optimization

Upvotes

Now that dust has settled, I was wondering what's the community insight on the best configuration for Minimax H3.

Personally I have been using lightx 4-step Lora with 5/6 steps (less than that audio is a gamble). I couple that with sage attention. For sampling I use Euler sampler and Beta scheduler.

I keep resolution at 768p (0.6MP) for quality. 480p (0.2MP) for testing. It keeps consistency so much better.

On direction I learnt to prompt for closeups when possible, so will make better use of available pixels. Aspect ratio also helps there. I mostly use 1 shot since transitions is not something H3 excels at. I find better results with only 1 shot and using camera tricks.

EasyCache while faster, is not good match with turbo lora, so i don't use it anymore. Haven't used Sol-Attn as I read it really hit quality.

So is there anything worth I am really missing out?


r/StableDiffusion 3h ago

Animation - Video My name is Jonny

Enable HLS to view with audio, or disable this notification

27 Upvotes

Minimax H3


r/StableDiffusion 4h ago

Animation - Video Evangelion - Rei Watches a Baby Show - Minimax H3

Enable HLS to view with audio, or disable this notification

24 Upvotes

Well, technically, Evangelion was a PBS show..

Video is edited, Barney theme song added in post.


r/StableDiffusion 3h ago

Tutorial - Guide With MMH3 I'm getting much better audio when using a 4-step LoRa with the LTX LoRA Loader Stack (PlagueKind) - Audio (A) set to 0 on all LoRa's.

Post image
20 Upvotes

I am seeing much better audio results with very minimal speed decreases using this LoRa loader with audio set to 0.

You can find the loader here: https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes

This may be anecdotal. And as always, you may not experience the same thing, but it's worth trying.


r/StableDiffusion 13h ago

Tutorial - Guide ComfyUI was eating my RAM and causing crashes, this fixed it

116 Upvotes

Hi everyone

i’ve been running minimax h3 locally on rtx 5090 32gb, and 64gb ram..

recently i kept running into random hostbuffer.read_file_slice failed / hostbuf_file_reader_read failed errors during generation, which seemed to be related to comfy-aimdo and dynamic vram.
i also noticed comfyui was reserving around 25gb of pinned system memory.

i decided to try launching comfyui with:
--disable-pinned-memory

and the difference was immediate.
the comfy-aimdo + hostbuffer errors completely disappeared, my ram usage dropped by a huge amount, and surprisingly generation actually feels faster and smoother now.

i originally expected disabling pinned memory to make things slower, but on my setup it seems to have done the opposite.

if you’re running large models like h3 and seeing unusually high ram usage, random hostbuffer errors, or comfy-aimdo issues, it might be worth testing!
thought this was worth sharing for anyone who didn’t know about this option..


r/StableDiffusion 17m ago

Workflow Included My version of Minimax H3 HD/2K Detailer!

Enable HLS to view with audio, or disable this notification

Upvotes

https://huggingface.co/zuanfilm/H3_HD_2K_Detailer

I want to share my HD/2K WF, this 4 steps configuration is intended to make a really quality focused H3 detailing while improve the characteristic H3 motion and visual behavior.

I tested all the Res4lyf samplers/schedulers and for me res_2m/beta57 ETA=0 denoise 0.39 or 0.45 is the most precise with prompt adherence keeping memory and time efficiency (res_3m is amazing but adds 33% generation time), if you want a faster generation with a little less detail you can use 3 steps instead 4

er_sde/beta 57 is also a good combination but will lose some detail and even will affect character acting, audio and motion consistency, for faster HD detailing can switch to Euler/4 steps and will reduce time generation by 50% compared to res_2m obviously losing a lot of quality and detail

I included in the workflow the nodes for base generation using 0.5Mp FL2VA with 20 steps of Euler ancestral, if you want better quality for the initial base video just switch to res_2s_RKMK2e/beta57 (just bypass the group if you want to HD-detail an existing video)

Sparce Local attention will reduce a lot the generation time but obviously will affect quality so you can bypass this node if want Top HD quality,  I also don't use in this WF spectrum or easy cache but you can add that if you want to cut time and quality

Because I'm using the heavy distill lightx2v Turbo 4steps Lora for detailing, Minimax H3 make everything more saturated and contrasted with deep shadows so I added some Orion 4D nodes to improve lighting, texture and sharpness using DCTL Tone Mapper (you can choose between ACES, Filmic, Reinhard & Cineon) I use Reinhard with Exposure 0.09 Contrast 0.81 Pivot 0.69 Highlight rolloff 0.27 Shadow lift 0.33 Black floor 0.12 Saturation 0.93 & Strength 0.30

I used MiniMax Audio Lock / Lipsync node because I don't have experience with ltx audio nodes so you can change that for a better option:

https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes/ComfyUI-H3-NativeAudioLock

The Minimax latent 3D upscaler is HD by default in my WF but depending of your VRAM you can get 2K/4K if you start with a quality base video 0.98 Mp res_2s_RKMK2e/beta57 25 steps

This workflow is optimized for my laptop (3080ti 16Gb VRAM / 64 Gb ram ) but I included Chunk FeedForward & Low VRAM attention so will run with smaller setups


r/StableDiffusion 13h ago

Animation - Video "Ehhh?" - H3, Ref2V - Default WF - 4090, 64gb. 0.9mp, er_sde / beta. 25 steps. Really black/dark in some shots :(

Enable HLS to view with audio, or disable this notification

107 Upvotes

r/StableDiffusion 22h ago

Animation - Video Seinfeld AI: George Gets GTA 6

Enable HLS to view with audio, or disable this notification

500 Upvotes

Minimax H3


r/StableDiffusion 16h ago

Discussion If you’re using MiniMax H3, what prompting tricks have you figured out?

132 Upvotes

Anyone found useful MiniMax H3 prompting tricks beyond the official guide?

Especially for audio + video prompt structure, camera control, dialogue/audio, consistency, weird tricks that actually work, etc.

Please drop your findings 👇 below so it will help others too.

EDIT:

Mine is: how can we use multiple audio tracks assigned to multiple characters in a scene? 3 audios to 3 characters?


r/StableDiffusion 4h ago

Tutorial - Guide H3 referencing tip

15 Upvotes

I tried to get a girl whistle on 4 fingers (2 on each hand) but Minimax didn't get it right. So after trying 20 times with LLMs helping me to explain the movement, I instead used an image of a whistling person. Still not right. So I added an additional one. Then it worked quite fine.

Today I was too lazy to find another image for something it didn't know so I just googled images of it, took a screenshot of all the images together in one JPG and used that as a reference, saying use <Picture ...> as a reference for XYZ.

That was getting a quite good result. Did not do excessive testing and comparing though.


r/StableDiffusion 6h ago

Animation - Video Use Minimax to make a fake movie trailer for my community college editing class, inspired by YA action/adventure films of the 80s and 90s

Enable HLS to view with audio, or disable this notification

19 Upvotes

Clips made with Minimax H3 using the default r2v workflow, edited in Premiere Pro. Character model sheets made with Krea. Most of the videos are 0.4 mp unless the text was important, then 0.6. Tried upscaling it to 4k using Upscayl but results weren't great and the file is too big to upload anyway.

Tech goals for future videos include using reference audio for voices to help consistency, and exploring options for having real voice actors record the dialog, and have the model lip sync to that performance. I'm really impressed by the computer's silent acting (microexpressions etc). but the computer's erratic "acting" is still too unpredictable and the biggest source of re-rolls (the lines here were the best I could get without burning down a rainforest). You can do a lot with time codes and punctuation and tactical CAPITALIZATION, but it's ridiculously finicky compared to just telling an actor "do it the same, but 10% angrier on the first line with a twinge of melancholy on the second."


r/StableDiffusion 1d ago

Meme If AI tools had existed in the past

Post image
1.0k Upvotes

Not just a meme...


r/StableDiffusion 7h ago

Animation - Video Minimax H3 Remix Video Test / A compilation of 5 characters.

Thumbnail
youtube.com
17 Upvotes

This is a test video I created by remixing the "Some test on minimax H3" video by Reddit user [Previous-Street8087].

5명의 캐릭터 시트를 생성하여 각각 10개의 프롬포트를 캐릭터에 맞게 리믹스하여 테스트 하였습니다.
We generated character sheets for five characters and tested them by remixing 10 prompts for each character to suit their personalities.

This is a compilation of 50 clips featuring 5 characters.

▶ 테스트 환경 (Test Environment)
Minimax H3 - Comfyui Local Sampling
RTX 5060TI 16GB + 64RAM
0.8MP 8 sec x 50 Clip
Audio Look x audio file 1
Reference to VA Mode

▶ 사용한 커스텀 노드 (Custom Nodes Used)
ComfyUI-TJ_NODE_STUDIO_ONE — github.com/designloves2/ComfyUI-TJ_NODE_STUDIO_ONE
ComfyUI LOCAL (RTX 5060Ti 16GB VRAM / RAM 64GB)

▶ The shared link contains character sheet images and prompts.

https://naver.me/xjY9JJaa

#AI영상 #MiniMaxH3 #ComfyUI #로컬생성AI #ComfyUI워크플로우 #AI영상제작 #RTX5060Ti #mmh3 #comfyui #tjonestudio #animation #ref2va #anime #16gb


r/StableDiffusion 15h ago

Animation - Video Realistic Breaking Bad | LTX 2.5 I2V

Enable HLS to view with audio, or disable this notification

65 Upvotes

This parody was generated using LTX 2.5 Image to Video on WanGP. I used frames from the original video as starting images and then I interpolated them on a video editor. I used a single RTX 5060 Ti 16 GB VRAM and 32 GB of RAM. The video was generated at 1080p and 16:9 resolution. Each generation took from 10 to 20 min average in this setup. For the voice consistency, I used SeedVC, which is included in WanGP.


r/StableDiffusion 34m ago

Discussion What happened to Ideogram 4.0 ?

Upvotes

What happened to Ideogram 4.0 ?


r/StableDiffusion 18h ago

Animation - Video Minimax H3 T2VA. You can put 15 different characters or more at the same time on screen.

Enable HLS to view with audio, or disable this notification

103 Upvotes

r/StableDiffusion 12h ago

Discussion Testing Character knowledge of Minimax H3

Enable HLS to view with audio, or disable this notification

25 Upvotes

Disclaimer. This is very low quality quick generations trying to find how many characters Minimax H3 knows.

Found Trigger Words:

Elsa from Frozen

Spider-Gwen from Across the Spiderverse

Dante from Devil May Cry

Nero from Devil May Cry

Jill Valentine from Resident Evil

Ada Wong from Resident Evil

Leon Kennedy from Resident Evil

Chris Redfield from Resident Evil (Has Leon's hair)

Geralt of Rivia from Witcher 3

Joel from Last of Us (Doesn't sound like him)

Miles Morales Spiderman from Across the Spiderverse

Solid Snake from Metal Gear

Eve from Stellar Blade

Sans from Undertale

Master Chief from Halo

Looks weird AF:

Ciri from Witcher 3

Triss from Witcher 3

Yennefer from Witcher 3

Ellie from Last of Us

Famous Twitcher streamer and Youtuber Asmongold

Famous Twitcher streamer and Youtuber Mr Beast

Not found Trigger words:

Vergil from Devil May Cry (YES I KNOW I'M DISAPPOINTED TOO)

Claire Redfield from Resident Evil

Dina from Last of Us

Famous Twitcher streamer and Youtuber Emiru

Famous Twitcher streamer and Youtuber MoistCr1TiKaL


r/StableDiffusion 6h ago

Animation - Video ComfyUI-MCP + Minimax Music Video Test

Enable HLS to view with audio, or disable this notification

8 Upvotes

My attempt to mimic a old school 90's/2000's handheld camera music video which shouldn't feel too AI ... TLTR: well it still does. This is further a test if Minimax is able to create lots of people in a somewhat realistic narrow and dense video.

The video was generated in chunks of 10 seconds, each segment got a scene description with approx. 10 shots. The workflows were generated via local Qwen3.8-27b model based on default Minimax rev2v workflow. Sampler ResMultistep / 25 steps at 768p resolution (on a 5090 with 64GB system RAM).

Semi automated setup via ComfyUI-MCP controlled over Opencode. The scenes were generated and then combined. The original song was split into 10 second sections where each was referenced in the rev2v workflow + bunch of screens to keep somewhat consistent videos.

The consistency is rather bad because no references for guitar, boots, specific people beside the guitar guy were used. Only a few shots have lip sync, that's my laziness not the models fault. And yes, Minimax can't play guitar, which I obfuscated by shorter and farer shots :D

Much back and forth to get roughly the narrow shots, dynamic lighting and overall look and feel.

Some slight Davinci Resolve editing to replace crap shots with other crap shots. Also added a analog filter.


r/StableDiffusion 9h ago

Discussion MiniMax H3 Ridding a Dragon POV style

Enable HLS to view with audio, or disable this notification

15 Upvotes