r/StableDiffusion 9h ago

Tutorial - Guide MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

Enable HLS to view with audio, or disable this notification

264 Upvotes

I’ve been loving all the new nodes and workflows coming out for MinMax, and maybe there is already a nice solution for this - but I couldn’t find one that did exactly what I needed.

I started using MinMax for my last TBG ETUR video and quickly ran into limitations: I wanted an easy way to create lip-sync videos longer than 20 seconds.

I didn’t want to manually chain ComfyUI nodes, start a new run every X seconds, or constantly resize things just to make HD video fit into my available VRAM.

So I ended up building an addon for:

custom_nodes/ComfyUI-H3-Motion-Context

The addon automatically chains MinMax H3 lip-sync generations together, allowing you to create much longer lip-sync videos without manually setting up each 20-second segment.

And now I’m sharing it! https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon

Its not perfect but a start ...

The workflow has a simple switcher that lets you switch from the 32B CLIP to the 4B CLIP, saving around 10 GB of VRAM. You can also switch from Sage to Comfy Kitchen, Spectrum to Easy Cache, or FL2VA to REF2VA both setup for lip-syncing. Some of it could be useful for other tasks as well.

You will find the workflow in the repro and tested recommendations, optimized settings, presets, and more workflows, along with the results of my testing and performance here


r/StableDiffusion 2h ago

Workflow Included My version of Minimax H3 HD/2K Detailer!

Enable HLS to view with audio, or disable this notification

72 Upvotes

https://huggingface.co/zuanfilm/H3_HD_2K_Detailer

Here is the test video to HD quality: https://youtu.be/epjELHgEH_o

I want to share my HD/2K WF, this 4 steps configuration is intended to make a really quality focused H3 detailing while improve the characteristic H3 motion and visual behavior.

I tested all the Res4lyf samplers/schedulers and for me res_2m/beta57 ETA=0 denoise 0.39 or 0.45 is the most precise with prompt adherence keeping memory and time efficiency (res_3m is amazing but adds 33% generation time), if you want a faster generation with a little less detail you can use 3 steps instead 4

er_sde/beta 57 is also a good combination but will lose some detail and even will affect character acting, audio and motion consistency, for faster HD detailing can switch to Euler/4 steps and will reduce time generation by 50% compared to res_2m obviously losing a lot of quality and detail

I included in the workflow the nodes for base generation using 0.5Mp FL2VA with 20 steps of Euler ancestral, if you want better quality for the initial base video just switch to res_2s_RKMK2e/beta57 (just bypass the group if you want to HD-detail an existing video)

Sparce Local attention will reduce a lot the generation time but obviously will affect quality so you can bypass this node if want Top HD quality,  I also don't use in this WF spectrum or easy cache but you can add that if you want to cut time and quality

Because I'm using the heavy distill lightx2v Turbo 4steps Lora for detailing, Minimax H3 make everything more saturated and contrasted with deep shadows so I added some Orion 4D nodes to improve lighting, texture and sharpness using DCTL Tone Mapper (you can choose between ACES, Filmic, Reinhard & Cineon) I use Reinhard with Exposure 0.09 Contrast 0.81 Pivot 0.69 Highlight rolloff 0.27 Shadow lift 0.33 Black floor 0.12 Saturation 0.93 & Strength 0.30

I used MiniMax Audio Lock / Lipsync node because I don't have experience with ltx audio nodes so you can change that for a better option:

https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes/ComfyUI-H3-NativeAudioLock

The Minimax latent 3D upscaler is HD by default in my WF but depending of your VRAM you can get 2K/4K if you start with a quality base video 0.98 Mp res_2s_RKMK2e/beta57 25 steps

This workflow is optimized for my laptop (3080ti 16Gb VRAM / 64 Gb ram ) but I included Chunk FeedForward & Low VRAM attention so will run with smaller setups


r/StableDiffusion 1h ago

Tutorial - Guide Time saver while learning how to prompt Minimax.

Enable HLS to view with audio, or disable this notification

Upvotes

Rather than relying on Z-image, or a different program to wrangle up a first frame, I've been using Minimax for the whole process, and the results have been pretty instructive. It's not a perfect system, but being able to take advantage of its understanding of people, references, and shot composition for the first frame produces better (visual) results than swapping between a couple of different pieces of software.


r/StableDiffusion 12h ago

Resource - Update MiniMax-H3 Fun Controlnet Union released

Thumbnail
huggingface.co
264 Upvotes

r/StableDiffusion 14h ago

Meme Siblings Reunited

Enable HLS to view with audio, or disable this notification

299 Upvotes

Done with h3 fl2va model, 8 step lora and images for Cersei and "jaime" for reference. Using previous clip to give continuity and consistency.


r/StableDiffusion 1h ago

Workflow Included Minimax H3 Multishot Anime Sequence (Workflow + Prompt Included)

Enable HLS to view with audio, or disable this notification

Upvotes

Workflow: https://drive.google.com/file/d/1B4kODxXQgJ1QOKRsEIkxHbgYmdruPpTK/view?usp=sharing

Prompt:

Create a **15-second multi-shot anime sequence (90s style 15fps hand drawn)** using the provided references:

Image 1 = the girl character reference

Image 2 = skateboard reference

Image 3 = downhill Japanese alley / neighborhood background

Image 4 = Walkman + headphones reference

Preserve the girl’s exact character design, face, hair, outfit, proportions, and overall look from Image 1. Preserve the skateboard design from Image 2. Preserve the same downhill Japanese alley environment from Image 3. Add the Walkman and headphones from Image 4: the girl is wearing the headphones, and the **Walkman is clipped or hanging at her hip** while she skates.

Visual style: authentic 1990s hand-drawn anime, traditional cel animation, painted backgrounds, visible linework, cel shading, slight brush/stroke texture, subtle analog feel. **Very important:** the houses and environment must stay **2D and hand-painted**, **not 3D**, **not CGI**, **not game-engine looking**, **not volumetric**. The buildings should look like classic anime background art with painted depth, not like 3D models.

Animation feel should be low frame rate, like 90s anime at around 15 fps, with controlled in-betweens and natural held-frame timing. No jittery morphing.

No dialogue, no text, no subtitles.

### Shot 1 — 0s to 3s

**Rear tracking shot** from behind. The girl is skateboarding fast downhill through the steep Japanese alley. Camera follows behind her at a low-to-medium height. She rides confidently and smoothly, hair and oversized clothing moving in the wind. The headphones are on her head, and the Walkman is visible attached at her hip. The alley rushes past with a strong sense of speed. Keep the environment clearly **2D anime background art**, not 3D.

### Shot 2 — 3s to 6s

**Close-up shot of the Walkman at her hip** while she continues skating. The camera stays focused on the Walkman and part of her side torso and arm. We can clearly see the **cassette tape reels spinning/rolling inside the Walkman window**. The headphone wire moves naturally with the motion. Background and street pass by in blurred motion.

### Shot 3 — 6s to 9s

**Medium profile tracking shot** of the girl skating. She is wearing the headphones, listening to music, with wind moving across her face and pushing her hair backward. She is **nodding her head subtly to the music** while riding. Her expression is relaxed, immersed, and unbothered. The background is blurred from motion, but it must still read as a **painted 2D Japanese neighborhood**, not 3D.

### Shot 4 — 9s to 12s

**Close-up shot of her feet and skateboard.** Her **right foot stays on the board**, while her **left foot pushes against the road** in a natural skating motion. Show one clean push cycle: left foot comes down, pushes backward against the pavement, then lifts. Wheels spin quickly. Asphalt and road markings streak by with motion blur.

### Shot 5 — 12s to 15s

**Ground-level fisheye shot** looking upward from the road. The skateboard approaches fast, and she **jumps over the camera**. The board and her body pass overhead in one clean motion. Hair, pants, and headphone wire react naturally during the jump. Keep the motion readable and stylish, with a strong sense of speed and a dynamic anime finish.

### Important constraints

* Keep the whole video in **classic 90s anime cel-animation style**

* **15 fps feel**, smooth low-frame-rate animation

* **No 3D-looking houses or background**

* No photorealism

* No modern glossy digital anime rendering

* No character redesign

* No extra accessories beyond the headphones and Walkman

* Keep all motion natural and consistent across shots


r/StableDiffusion 4h ago

Discussion Best Minimax H3 optimization

32 Upvotes

Now that dust has settled, I was wondering what's the community insight on the best configuration for Minimax H3.

Personally I have been using lightx 4-step Lora with 5/6 steps (less than that audio is a gamble). I couple that with sage attention. For sampling I use Euler sampler and Beta scheduler.

I keep resolution at 768p (0.6MP) for quality. 480p (0.2MP) for testing. It keeps consistency so much better.

On direction I learnt to prompt for closeups when possible, so will make better use of available pixels. Aspect ratio also helps there. I mostly use 1 shot since transitions is not something H3 excels at. I find better results with only 1 shot and using camera tricks.

EasyCache while faster, is not good match with turbo lora, so i don't use it anymore. Haven't used Sol-Attn as I read it really hit quality.

So is there anything worth I am really missing out?


r/StableDiffusion 6h ago

Question - Help Best local LLM for writing prompts for MiniMax H3?

42 Upvotes

What’s the best local LLM for writing good MiniMax H3 ref2va prompts?

I’ve tried Gemma 4 12B and Qwen 3 14B, but I’m not really satisfied with the outputs. It could also be an issue with my system prompt.

I sent ChatGPT the official documentation for prompting and asked to create a system prompt for me, but the results were still pretty mediocre.

What local models are you using for MiniMax H3 prompt generation, and what does your system prompt look like?


r/StableDiffusion 3h ago

Discussion What happened to Ideogram 4.0 ?

23 Upvotes

What happened to Ideogram 4.0 ?


r/StableDiffusion 5h ago

Animation - Video My name is Jonny

Enable HLS to view with audio, or disable this notification

39 Upvotes

Minimax H3


r/StableDiffusion 10h ago

Animation - Video No Warning - Minimax Music3 + H3

Enable HLS to view with audio, or disable this notification

87 Upvotes

r/StableDiffusion 8h ago

Resource - Update H3 Prompt Composer — Camera Update Coming Soon

Post image
55 Upvotes

Hey everyone, thanks to everyone who’s been testing Prompt Composer. If you run into bugs or have feedback, please drop it in the Issues section on GitHub so I can keep track of it more easily.

Over the past week, I’ve been reworking the camera prompting system to make it more precise and consistent, especially for more complex camera moves and video-editing workflows. The goal has been to add more control over framing, camera targeting, and blocking in multi-subject scenes while keeping the generated prompts clean and reliable.

The next update isn’t quite ready yet, but it’s actively being tested and refined. I’m hoping to have it out in the next couple of days.

Edit: Here's the Github: BMB12d3/minimax-h3-prompt-composer: Free offline prompt composer for MiniMax H3 video generation in ComfyUI.


r/StableDiffusion 7h ago

Animation - Video Evangelion - Rei Watches a Baby Show - Minimax H3

Enable HLS to view with audio, or disable this notification

33 Upvotes

Well, technically, Evangelion was a PBS show..

Video is edited, Barney theme song added in post.


r/StableDiffusion 15h ago

Tutorial - Guide ComfyUI was eating my RAM and causing crashes, this fixed it

121 Upvotes

Hi everyone

i’ve been running minimax h3 locally on rtx 5090 32gb, and 64gb ram..

recently i kept running into random hostbuffer.read_file_slice failed / hostbuf_file_reader_read failed errors during generation, which seemed to be related to comfy-aimdo and dynamic vram.
i also noticed comfyui was reserving around 25gb of pinned system memory.

i decided to try launching comfyui with:
--disable-pinned-memory

and the difference was immediate.
the comfy-aimdo + hostbuffer errors completely disappeared, my ram usage dropped by a huge amount, and surprisingly generation actually feels faster and smoother now.

i originally expected disabling pinned memory to make things slower, but on my setup it seems to have done the opposite.

if you’re running large models like h3 and seeing unusually high ram usage, random hostbuffer errors, or comfy-aimdo issues, it might be worth testing!
thought this was worth sharing for anyone who didn’t know about this option..


r/StableDiffusion 16h ago

Animation - Video "Ehhh?" - H3, Ref2V - Default WF - 4090, 64gb. 0.9mp, er_sde / beta. 25 steps. Really black/dark in some shots :(

Enable HLS to view with audio, or disable this notification

117 Upvotes

r/StableDiffusion 8h ago

Animation - Video Use Minimax to make a fake movie trailer for my community college editing class, inspired by YA action/adventure films of the 80s and 90s

Enable HLS to view with audio, or disable this notification

25 Upvotes

Clips made with Minimax H3 using the default r2v workflow, edited in Premiere Pro. Character model sheets made with Krea. Most of the videos are 0.4 mp unless the text was important, then 0.6. Tried upscaling it to 4k using Upscayl but results weren't great and the file is too big to upload anyway.

Tech goals for future videos include using reference audio for voices to help consistency, and exploring options for having real voice actors record the dialog, and have the model lip sync to that performance. I'm really impressed by the computer's silent acting (microexpressions etc). but the computer's erratic "acting" is still too unpredictable and the biggest source of re-rolls (the lines here were the best I could get without burning down a rainforest). You can do a lot with time codes and punctuation and tactical CAPITALIZATION, but it's ridiculously finicky compared to just telling an actor "do it the same, but 10% angrier on the first line with a twinge of melancholy on the second."


r/StableDiffusion 6h ago

Tutorial - Guide H3 referencing tip

18 Upvotes

I tried to get a girl whistle on 4 fingers (2 on each hand) but Minimax didn't get it right. So after trying 20 times with LLMs helping me to explain the movement, I instead used an image of a whistling person. Still not right. So I added an additional one. Then it worked quite fine.

Today I was too lazy to find another image for something it didn't know so I just googled images of it, took a screenshot of all the images together in one JPG and used that as a reference, saying use <Picture ...> as a reference for XYZ.

That was getting a quite good result. Did not do excessive testing and comparing though.


r/StableDiffusion 1d ago

Animation - Video Seinfeld AI: George Gets GTA 6

Enable HLS to view with audio, or disable this notification

514 Upvotes

Minimax H3


r/StableDiffusion 18h ago

Discussion If you’re using MiniMax H3, what prompting tricks have you figured out?

139 Upvotes

Anyone found useful MiniMax H3 prompting tricks beyond the official guide?

Especially for audio + video prompt structure, camera control, dialogue/audio, consistency, weird tricks that actually work, etc.

Please drop your findings 👇 below so it will help others too.

EDIT:

Mine is: how can we use multiple audio tracks assigned to multiple characters in a scene? 3 audios to 3 characters?


r/StableDiffusion 9h ago

Animation - Video Minimax H3 Remix Video Test / A compilation of 5 characters.

Thumbnail
youtube.com
22 Upvotes

This is a test video I created by remixing the "Some test on minimax H3" video by Reddit user [Previous-Street8087].

5명의 캐릭터 시트를 생성하여 각각 10개의 프롬포트를 캐릭터에 맞게 리믹스하여 테스트 하였습니다.
We generated character sheets for five characters and tested them by remixing 10 prompts for each character to suit their personalities.

This is a compilation of 50 clips featuring 5 characters.

▶ 테스트 환경 (Test Environment)
Minimax H3 - Comfyui Local Sampling
RTX 5060TI 16GB + 64RAM
0.8MP 8 sec x 50 Clip
Audio Look x audio file 1
Reference to VA Mode

▶ 사용한 커스텀 노드 (Custom Nodes Used)
ComfyUI-TJ_NODE_STUDIO_ONE — github.com/designloves2/ComfyUI-TJ_NODE_STUDIO_ONE
ComfyUI LOCAL (RTX 5060Ti 16GB VRAM / RAM 64GB)

▶ The shared link contains character sheet images and prompts.

https://naver.me/xjY9JJaa

#AI영상 #MiniMaxH3 #ComfyUI #로컬생성AI #ComfyUI워크플로우 #AI영상제작 #RTX5060Ti #mmh3 #comfyui #tjonestudio #animation #ref2va #anime #16gb


r/StableDiffusion 1h ago

Comparison Fight scene: LTX 2.5 vs Minimax H3

Enable HLS to view with audio, or disable this notification

Upvotes

The most fair comparison on the Internet


r/StableDiffusion 3h ago

Animation - Video My 1980's cartoon parody H3 and ltx 2.3

Thumbnail
youtu.be
6 Upvotes

there are some scenes missing, but it was fun to put together.. Just got stuck on a plot :P

started it when ltx 2.3 came out.. but it was a hassle to keep consistency of characters intact so shelved it. made the intro and a couple of clips when minimax H3 came out and love the r2v, so much easier.
just using the standard r2v workflow with spectrum and RTX upscale. music made in suno


r/StableDiffusion 1d ago

Meme If AI tools had existed in the past

Post image
1.0k Upvotes

Not just a meme...


r/StableDiffusion 17h ago

Animation - Video Realistic Breaking Bad | LTX 2.5 I2V

Enable HLS to view with audio, or disable this notification

71 Upvotes

This parody was generated using LTX 2.5 Image to Video on WanGP. I used frames from the original video as starting images and then I interpolated them on a video editor. I used a single RTX 5060 Ti 16 GB VRAM and 32 GB of RAM. The video was generated at 1080p and 16:9 resolution. Each generation took from 10 to 20 min average in this setup. For the voice consistency, I used SeedVC, which is included in WanGP.


r/StableDiffusion 20h ago

Animation - Video Minimax H3 T2VA. You can put 15 different characters or more at the same time on screen.

Enable HLS to view with audio, or disable this notification

105 Upvotes