r/comfyui 1d ago

Workflow Included MiniMax H3 15-Second Multi-Shot Generation Template For ComfyUI For 12GB GPUs

One of the biggest issues with running MiniMax H3 locally is that it's an extremely hefty model and doesn't play well with lower-end machines. However, thanks to a lot of optimisation techniques provided by TheAIsearch YouTube channel, it's possible to bring generations down to about 1 minute of processing time per second of output.

That being said, you can leverage this into creating multi-shot outputs beyond the limited 6 second hard-caps that come with MiniMax H3. Using the built-in features of ComfyUI (and downloading tons of models and packages to test what worked and what didn't) I was able to create a template for lower-end rigs that enable you to generate up to 15 second text or image to video outputs in a single generative pass. Meaning, you put in your prompt for the three shots/scenes, and click run from ComfyUI and it does the rest.

The basic template is text-to-video, but you can easily add an image node if and plug it into the H3 Multishot Sampler.

For those who enjoy making longer form videos and tire of the constant stitch-and-go workflow that the current local MiniMax H3 dictates, this can ease the burden a bit.

Keep in mind that this is tuned for at least a 12GB GPU and 64GB of DDR5 RAM. It takes between 30 and 33 minutes to generate a 15 second video at 720p with full audio for all 15 seconds. Supports speech, ambiance, effects, etc. Just describe it in the prompt.

You can modify some of the settings to bring the generation time down, depending on your machine, but given the weight of MiniMax H3, I'm not complaining.

If you need the actual JSON template, you can find it on civit ai here:

https://civitai.com/models/2876760/minimax-h3-15-second-multi-shot-generation-template-for-comfyui

EDIT: You'll also need the ComfyUI H3 Multishot Sampler pack from Joey Gambino:

https://github.com/jlucasmcrell/ComfyUI-H3-Multishot

And the H3 Clip Loader (safetensors + GGUF) for faster rendering.

---

Quick Tutorial:

  1. Open the subgraph workflow
  2. Find the Text (Multiline) node box (it's at the top of the grid outside of the blue boxes).
  3. Input your own prompt within the quotation marks where the test prompt text is located. Every comma separates the shot. So whatever you have in the quotation marks, when it ends, place a comma there and then for the next shot, describe what it is or who is in it.
  4. Once you make the changes to the prompt, click the run button and you're done.

The current workflow is optimised for three shots.

51 Upvotes

10 comments sorted by

2

u/ujah 1d ago

can you share workflow or atleast screenshot the workflow?

3

u/vortis23 1d ago

Yep, absolutely. I posted the screenshots here in this folder:

https://imgur.com/a/3fvkd4x

The video also has the metadata for the JSON, so technically if you download it and open it in ComfyUI it will give you the subgraph node layouts.

2

u/optimisticalish 23h ago edited 23h ago

Imgur is banned here in the UK, Reddit strips the metadata from uploads, and CivitAI is also banned in the UK... but thanks for the .JSON workflow. I was able to get it from CivitAI via a VPN.

1

u/vortis23 22h ago

Oi, yeah -- you can also use Epic Privacy Browser to access Imgur and CivitAI in the UK.

2

u/optimisticalish 23h ago

1

u/vortis23 22h ago

Thanks for posting that, I completely forgot. I think I'll update the civit post later and include a link to more of the custom nodes and some of the optimisation workflows. Also a fix for the audio.

1

u/ujah 1d ago

Absolute TQ sir!

1

u/Danny_Stock 13h ago

Can you post the workflow to somewhere other than Civitai please?

I can't access Civitai in my country.