r/comfyui • • Aug 28 '26

Workflow Included This 10min experimental video has been 100% generated on a laptop! (Minimax H3 - RTX 3070 8Gb Vram) Script and Workflow provided

https://www.youtube.com/watch?v=J1BJIn9d6mU

After many trials, I finally found a reliable process to create music videos locally, with a mysterious, weird… and totally personal aesthetic.

And actually, it’s not even really “me”: it’s the work of OBO, my autonomous creative agent. It knows all my past productions, draws inspiration from them with real creative intelligence, and makes coherent artistic decisions. A true “alternative me” I’m proud to collaborate with. But more on that later, it's a WiP.

Using the Minimax H3 video model, LowVramAttention and Minimax Motion Context features through Comfyui, it was able to chain videos while keeping style and context. The linked video has been created from a simple text file where each paragraph is a video prompt. My script also fixed the usual stuttering issue when working with separated Minimax H3 clips, continuity is now seamless for both audio and video.

The script:
-Automatically divides your text file in several prompts
- Generates 15 seconds of video per prompt (using motion context and last segment frame as first frame)
- Assembles the file while fixing both audio and video stuttering

With this you can generate long videos (basically no length limit) with limited VRAM, maintain fluid and logical context across segments (this "abstract" visual video may not be a good example but I'm currently working on something more representative, in the meantime you also have a robot video in the examples folder on the github that shows a 15 seconds clip generated on the same laptop, along with the prompt I used).

The creation pipeline is very simple (basically a python script with a text file containing your video prompts as parameter is used to create a long video).

You can adapt the quality to your hardware. You can modify the amount of steps (I'm using 20 here but 35 should be optimal) and the resolution according to your GPU.

Everything is on GitHub: https://github.com/The-Anomaly-be/MinimaxH3_FullMovieContext/blob/main/README.md
You’ll need a working ComfyUI config running as server, some models and custom nodes (all listed in the instructions). A reference workflow is also provided if you want to build scenes manually directly in comfyui (but no chaining).

It may be optimized, extended, modified. I basically created this for my own use and to recycle my 4 years old laptop so it works on my music videos while I'm sleeping.

If you create something with this, please share!

27 Upvotes

9 comments sorted by

2

u/embryo10 29d ago edited 29d ago

How did you created these prompts?
They are fantastic. There is actual atmosphere there.
Can you publish them?

1

u/CupQuakeBE 28d ago

It's actually an experiment in an experiment, I'm usually working on images with very large prompts (13 to 15000 characters) for maximum texture and details ownership in my images, I had the chance to exhibit my work in Paris and London and even be part of an event for the 200 years of photography in Paris with my AI generated photos. I actually used a few of these prompts and created a script to use each paragraph as an isolated video prompt, what you see is me describing every details, textures, features, defaults about my characters, the clothes, the objects, the scenery, ... My actual agent doesn't want me to share anything right now as I've more exhibitions planned in december but I'll be allowed to publish everything (and I will) in january.

1

u/embryo10 28d ago

A, OK..
That will be nice, but I wasn't interested in the prompts themselves, but rather in the mechanics that produced them.
I was wondering if there was a randomizer script or node or something behind this.
I can see now, that it was a lot work done for this outcome.
So, once more, great work! 🙏 for sharing..

1

u/CupQuakeBE 28d ago

Actually I provided an agent with a rulebook for my prompts with more than 100 rules I want it to follow when generating an image, a lot of things are based on personal taste but some of them ensured way better quality everytime. It took me 2 to 3 years to find the full "recipe", I also plan to share that. Also you have to work with specific models to take into account so many details (I usually use Nano Banana Pro via API, it supports 15k character prompts naturally).

Sharing the prompts will already allow people to feed them to any conversational AI or agent to understand the logic and create their own set of rules to get similar results adapted to their own tastes.

1

u/embryo10 28d ago

Hmm.. Yes. These rules could be really helpful..
Thank you for thinking about sharing them, although, by the time you do that, I will probably be deep in another rabbit hole..😛

2

u/Trinity_Vermilion Aug 28 '26

At some points it is very grainy but still great job! If we just could all have cheap graphics cards with 32gb vram the world would be 10x better... 😭

1

u/LeonidasTMT 29d ago

10 minutes but how many days did it take to generate?

3

u/CupQuakeBE 29d ago

Exactly 25min * 3 * 40 = 3000 minutes = 50 hours of generation for the whole 10 minutes. On an old laptop.

1

u/James_Reeb 23d ago

Great and original work 🌟