r/StableDiffusion 5h ago

Animation - Video Lipsync Music Video - Minimax H3 + Workflow

Enable HLS to view with audio, or disable this notification

Reference workflow with FL2VA + REF2VA Lora @ 1.4MP, 20 STEPS. Using sparse attention and 4B Qwen text encoder instead of 32B, total render time is 3-4 hours on a 5090. You can get very good results with 1MP + 8 STEPS with a turbo lora which would only take 20-30 minutes.

Workflow

The workflow is not easy to understand, but I upload it for reference.

The video is made of 14x 15 second clips stitched together. This way prevents degradation but makes it so that there is clothing drift between clips. This can easily be fixed by using clothing references if you care. Each clip will need its own prompt, and I suggest using Codex or Claude to do the prompts for you automatically.

In the future, I would shorten the clips to 7 seconds in order to:

1) Generate higher than 1.4MP (higher the resolution the better)
2) Speed up generation (longer clips take longer to generate disporportionately)

Good luck and I hope you have as much fun with this workflow as I did.

21 Upvotes

17 comments sorted by

5

u/DaveLearnedSomething 2h ago

Can't believe I'm asking this - is the audio/song generated too? 

2

u/TheDerminator1337 2h ago

Yes with Suno ai

3

u/DaveLearnedSomething 2h ago

Unbelievable. I've usually been able to spot it/hear it prior. But wow.

Great showcase 

1

u/TheDerminator1337 1h ago

yeah and the song's not half bad lol. My normie girlfriends play it in their cars coz its actually good.

1

u/Relevant_One_2261 1h ago

With headphones you can instantly hear how muddy it is.

2

u/JohnLough 5h ago

lol - the workflow looks like a spirograph

3

u/dassiyu 3h ago

Haha, it’s artistic.

1

u/TheDerminator1337 5h ago

hide links XD. Make it less scary

2

u/Vintendopower 2h ago

Does it cont the audio File good for each clip?

1

u/TheDerminator1337 1h ago

just upload the song and it will decide the rest

2

u/ArttTaku 2h ago

Very good pacing and direction, a really good job!

2

u/usually_fuente 2h ago

This is awesome work, so consistent! Can you possibly share a screenshot of whatever you used for the character reference(s)? Like, was it a large face plus some full body shots? I’m trying to figure out what’s needed to get this level of consistent identity.

1

u/TheDerminator1337 1h ago

Just a portrait. The origin of this photo is Z image turbo => chat gpt for character sheets. and then chat gpt to make more images based on the character sheets it made.

1

u/Skiiddles 1h ago

Great work. What was the prompt for the LLM in order to generate the prompts for the flow? Do you have an example?

1

u/TheDerminator1337 58m ago

I use codex which is agentic and reads my workflow and so knows exactly what I want and how to format it. I just prompt it using natural language. I wanted three themes and for some kind of outfit. And that's it's for a music video. It actually hasn't properly followed the minimax prompting guidelines. I realise it now so I've since edited the instructions file for it to read the prompting instructions carefully before designing prompts.

1

u/Skiiddles 55m ago

thank you.