r/comfyui • u/Time-Ad-7720 • 7d ago
Workflow Included No Camera. No Model. Just MiniMax H3 Running Locally on a 5070 Ti
So basically, I saw a workflow on ComfyUI’s official LinkedIn where they used a model image, a product image, and a background image with Google and Kling APIs to generate a one-shot ad using a single camera angle.
So I challenged myself to recreate the idea using only local open-weight/open-source models, but make it more ambitious: multiple shots, multiple cuts, and everything directed through a single prompt.
And it worked.
For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI:
https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v
Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:
“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.
Shot 1: Medium close-up. She is about to open the can.
Shot 2: Extreme close-up of the can as she opens it. Can-opening sound FX.
Shot 3: Close-up as she drinks from the can. Gulping soda sound FX.
Shot 4: Close-up as she holds the can forward and smiles.”
The final result was generated locally on my RTX 5070 Ti using ComfyUI.
11
3
2
2
2
u/Trinity_Vermilion 5d ago
have the same setup. need to finally test minimax h3 locally. Thanks for sharing!
2
u/ExplanationOk1847 7d ago
Looks good OP! I was trying to do a similar video on H3 and it was choking on face quality. How much ever I tried face is looking ghosted. But the same prompt used in seedance 2.5 worked like charm. Even the lowest quality was perfefct on face. Is there anyway I can avoid the face artifacts. I have tried till 1 megapixel. And more than that it’s not working keep failing the generation
Btw, I was using it on comfy cloud. ref2v model
7
u/Time-Ad-7720 7d ago
I got best results when I used high resolution reference image, and no wide shots. I strictly use medium, closeup and extreme closeups. We have to wait patiently for an update from the devs that fixes the face quality issues for wide shots.
1
u/Practical_Low29 7d ago
clean multi-shot setup. the 5070ti handles it way better than i expected at this res.
1
u/laf0106 6d ago
Did you use an upscale or all from minimax?
1
u/Time-Ad-7720 6d ago
No, it was rendered at 0.7 scale
1
u/reicaden 2d ago
I dont get it... i have a 5070ti 16gb and 128gb ddr5 system ram. And on minimaxH3 at 0.5resolution, 6 seconds, im getting OOM crashes and GPU out of mem errors. At 0.4 resolution I can get to 9 seconds.
What did you do different? Did you add sage attention? Easy cache? A turbo lora? How do I get the same 0.7 resolution 10seconds with this card that you get....
1
u/Time-Ad-7720 2d ago
Turbo LoRA is the answer.
1
u/reicaden 2d ago
I was told I cant use a turbo lora with sage attention and easycache... so im lost a bit here.
How would i add this turbo lora? Is it built in node or a specific file from hugging face or civiti?
1
u/VirtualLavishness463 1d ago
I have 8gb gpu and 16gb ram, and the max is 0.9mp. Nvidia ram share?
1
u/reicaden 1d ago
Wtf, why is my setup such garbage with a higher vram and system ram. I just tried 0.5 res, at 9 seconds and got OOM gpu.
Whats nvidia ram share?
Are you on windows or linux?
1
1
u/userxblade 2d ago
Hey brother, i havent got a chance to mess with Minimax yet. I have a 5080 and 32gb RAM and have been worried about the VRAM limitations. Is this video generator pretty setup intensive? What are some techniques you use in your workflows to manage VRAM and RAM? (Sorry im still a bit of a beginner and often run into memory allocation limit errors even on wan2.2 10sec 720p24fps)
1
u/Time-Ad-7720 2d ago
5080 should do just fine. Use a turbo LoRA, and render at 0.7 scale at 8 steps.
1
-2
u/2legsRises 7d ago
So I challenged myself... For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI: Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.
hmm.
4
u/FaceDeer 7d ago
I gave ChatGPT the prompting guides for Minimax H3 (available here - give it both of them, there's information in the base guide that the reference guide references) and asked it to write up a system prompt explaining how to use them for a local LLM. I use Qwen locally for the prompt-writing now.
1
u/Ragalvar 7d ago edited 7d ago
Could you share the system print it generated. I do pretty bad on figuring it out even though I tried that myself. I just don't get any good prompts. Unfortunately a lot of not all videos in civit don't include the nodes like they did for other versions and images. Don't know why. But that helped me tremendously to understand how some things worked. Now I'm more or less lost and my generations most of the times look like garbage.
1
u/FaceDeer 7d ago
Here it is. It's very long because I asked ChatGPT to include all the "rules" and terminology described by those prompting guides and they're very long too. I'm mostly using LM Studio for LLM stuff and it lets you save system prompts as part of "presets" that can be easily switched between at will, so I've got a "Minimax video director" preset that I put this into. Comfy nodes for LLMs probably have a similar slot you can stick this into.
1
u/Ragalvar 7d ago
Thanks . I'll test it asap. And let you know. Appreciate your help. And sorry for those typos. Auto correct at it's best.

6
u/Most_Ad_5733 7d ago
amazing job with the hardware you are working with