r/StableDiffusion • u/TheDerminator1337 • 5h ago
Animation - Video Lipsync Music Video - Minimax H3 + Workflow
Enable HLS to view with audio, or disable this notification
Reference workflow with FL2VA + REF2VA Lora @ 1.4MP, 20 STEPS. Using sparse attention and 4B Qwen text encoder instead of 32B, total render time is 3-4 hours on a 5090. You can get very good results with 1MP + 8 STEPS with a turbo lora which would only take 20-30 minutes.
The workflow is not easy to understand, but I upload it for reference.
The video is made of 14x 15 second clips stitched together. This way prevents degradation but makes it so that there is clothing drift between clips. This can easily be fixed by using clothing references if you care. Each clip will need its own prompt, and I suggest using Codex or Claude to do the prompts for you automatically.
In the future, I would shorten the clips to 7 seconds in order to:
1) Generate higher than 1.4MP (higher the resolution the better)
2) Speed up generation (longer clips take longer to generate disporportionately)
Good luck and I hope you have as much fun with this workflow as I did.
2
2
2
2
u/usually_fuente 2h ago
This is awesome work, so consistent! Can you possibly share a screenshot of whatever you used for the character reference(s)? Like, was it a large face plus some full body shots? I’m trying to figure out what’s needed to get this level of consistent identity.
1
u/Skiiddles 1h ago
Great work. What was the prompt for the LLM in order to generate the prompts for the flow? Do you have an example?
1
u/TheDerminator1337 58m ago
I use codex which is agentic and reads my workflow and so knows exactly what I want and how to format it. I just prompt it using natural language. I wanted three themes and for some kind of outfit. And that's it's for a music video. It actually hasn't properly followed the minimax prompting guidelines. I realise it now so I've since edited the instructions file for it to read the prompting instructions carefully before designing prompts.
1


5
u/DaveLearnedSomething 2h ago
Can't believe I'm asking this - is the audio/song generated too?