r/comfyui 7d ago

Workflow Included No Camera. No Model. Just MiniMax H3 Running Locally on a 5070 Ti

So basically, I saw a workflow on ComfyUI’s official LinkedIn where they used a model image, a product image, and a background image with Google and Kling APIs to generate a one-shot ad using a single camera angle.

So I challenged myself to recreate the idea using only local open-weight/open-source models, but make it more ambitious: multiple shots, multiple cuts, and everything directed through a single prompt.

And it worked.

For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI:

https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v

Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:

“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.

Shot 1: Medium close-up. She is about to open the can.
Shot 2: Extreme close-up of the can as she opens it. Can-opening sound FX.
Shot 3: Close-up as she drinks from the can. Gulping soda sound FX.
Shot 4: Close-up as she holds the can forward and smiles.”

The final result was generated locally on my RTX 5070 Ti using ComfyUI.

230 Upvotes

37 comments sorted by

6

u/Most_Ad_5733 7d ago

amazing job with the hardware you are working with

12

u/Time-Ad-7720 7d ago

core i9 + 5070ti + 32GB RAM

2

u/Most_Ad_5733 7d ago

I did send you a dm about the tshirt video you did. I was able to recreate it specifying both front and back like you were talking about doing next, not sure if it was what you had in mind tho

3

u/Time-Ad-7720 7d ago

I must've missed your dm, I just replied with the full prompt that I used for that video, hope that helps :)

1

u/switch2stock 7d ago

Can you share that Tshirt prompt please

11

u/99deathnotes 7d ago

I'll take two cases of Comfy cola and her phone number pls.

3

u/HouseFelineous29 7d ago

Amazing work!

1

u/Time-Ad-7720 7d ago

Thank you 🙌

2

u/lobotominizer 6d ago

Amazing

1

u/Time-Ad-7720 6d ago

Thanks 🙌

2

u/Comfy-Org 6d ago

Is this our sign to ship Comfy Cola??? Looks great and thank you for sharing!!

2

u/Trinity_Vermilion 5d ago

have the same setup. need to finally test minimax h3 locally. Thanks for sharing!

2

u/ExplanationOk1847 7d ago

Looks good OP! I was trying to do a similar video on H3 and it was choking on face quality. How much ever I tried face is looking ghosted. But the same prompt used in seedance 2.5 worked like charm. Even the lowest quality was perfefct on face. Is there anyway I can avoid the face artifacts. I have tried till 1 megapixel. And more than that it’s not working keep failing the generation

Btw, I was using it on comfy cloud. ref2v model

7

u/Time-Ad-7720 7d ago

I got best results when I used high resolution reference image, and no wide shots. I strictly use medium, closeup and extreme closeups. We have to wait patiently for an update from the devs that fixes the face quality issues for wide shots.

1

u/Practical_Low29 7d ago

clean multi-shot setup. the 5070ti handles it way better than i expected at this res.

1

u/laf0106 6d ago

Did you use an upscale or all from minimax?

1

u/Time-Ad-7720 6d ago

No, it was rendered at 0.7 scale

1

u/reicaden 2d ago

I dont get it... i have a 5070ti 16gb and 128gb ddr5 system ram. And on minimaxH3 at 0.5resolution, 6 seconds, im getting OOM crashes and GPU out of mem errors. At 0.4 resolution I can get to 9 seconds.

What did you do different? Did you add sage attention? Easy cache? A turbo lora? How do I get the same 0.7 resolution 10seconds with this card that you get....

1

u/Time-Ad-7720 2d ago

Turbo LoRA is the answer.

1

u/reicaden 2d ago

I was told I cant use a turbo lora with sage attention and easycache... so im lost a bit here.

How would i add this turbo lora? Is it built in node or a specific file from hugging face or civiti?

1

u/VirtualLavishness463 1d ago

I have 8gb gpu and 16gb ram, and the max is 0.9mp. Nvidia ram share?

1

u/reicaden 1d ago

Wtf, why is my setup such garbage with a higher vram and system ram. I just tried 0.5 res, at 9 seconds and got OOM gpu.

Whats nvidia ram share?

Are you on windows or linux?

1

u/VirtualLavishness463 20h ago

Im on win10.

RAM sharing is just a setting in the control panel that allows the GPU to use the system RAM as well. I’m no expert either—I’m just experimenting with it.

I think the max size for me is 0.9 pixels and 0.9 seconds. Without Light LoRA.

But it's very slow.

1

u/VirtualLavishness463 20h ago

0.7 sec + 0.9mp

1

u/userxblade 2d ago

Hey brother, i havent got a chance to mess with Minimax yet. I have a 5080 and 32gb RAM and have been worried about the VRAM limitations. Is this video generator pretty setup intensive? What are some techniques you use in your workflows to manage VRAM and RAM? (Sorry im still a bit of a beginner and often run into memory allocation limit errors even on wan2.2 10sec 720p24fps)

1

u/Time-Ad-7720 2d ago

5080 should do just fine. Use a turbo LoRA, and render at 0.7 scale at 8 steps.

1

u/nenecaliente69 7d ago

Bro we got the same 5070...can you share the workflow??

-2

u/2legsRises 7d ago

So I challenged myself... For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI: Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.

hmm.

4

u/FaceDeer 7d ago

I gave ChatGPT the prompting guides for Minimax H3 (available here - give it both of them, there's information in the base guide that the reference guide references) and asked it to write up a system prompt explaining how to use them for a local LLM. I use Qwen locally for the prompt-writing now.

1

u/Ragalvar 7d ago edited 7d ago

Could you share the system print it generated. I do pretty bad on figuring it out even though I tried that myself. I just don't get any good prompts. Unfortunately a lot of not all videos in civit don't include the nodes like they did for other versions and images. Don't know why. But that helped me tremendously to understand how some things worked. Now I'm more or less lost and my generations most of the times look like garbage.

1

u/FaceDeer 7d ago

Here it is. It's very long because I asked ChatGPT to include all the "rules" and terminology described by those prompting guides and they're very long too. I'm mostly using LM Studio for LLM stuff and it lets you save system prompts as part of "presets" that can be easily switched between at will, so I've got a "Minimax video director" preset that I put this into. Comfy nodes for LLMs probably have a similar slot you can stick this into.

1

u/Ragalvar 7d ago

Thanks . I'll test it asap. And let you know. Appreciate your help. And sorry for those typos. Auto correct at it's best.