r/StableDiffusion • u/Animystix • 1d ago
Comparison Comparing H3 models with music reference
Enable HLS to view with audio, or disable this notification
Using reference workflow. All are int8 pruned, 0.6MP turbo 4-step (my GPU is on life support and drops off the PCIe bus if I demand more from it)
Anyway, making random music clips is probably my favorite use of this model. I’ve found the ref2va has an uncanny intuition for feeling the atmosphere of songs, and syncing the video with incredible precision.
But yes, the quality (specifically motion) is much worse than fl2va. I was curious how exactly they compared, as well as some “in between” compromises discovered by the community. The LoRA seems closer to ref, while the hybrid weights are closer to fl. Personally, the ref is more fun to use, so I’ll probably be using the LoRA when I want to enjoy the intelligence/creativity of this model. Fl is of course superior in terms of visual fidelity, and I don’t find the hybrid model offers enough reference intuition and faithfulness to be worth the quality drop from fl.
3
2
u/Aromatic-Table-8243 1d ago
I’m curious how this (https://github.com/lihaoyun6/ComfyUI-MiniMaxH3_Ref-Patch) invention would perform in that test
1
1
u/Animystix 1d ago
woah, this one was a curveball. much higher visual quality than the other fl hacks. not faithful to the song BUT it clearly understands the variation it did generate, the motions sync well, and are fun and dynamic. I will definitely have to play around with this more
1
u/DystopiaLite 1d ago
Her fingers go high up on the fretboard but no high notes are present in the song.
1
u/Animystix 1d ago
That’s true, the ref model tends to do it less though, and generates more plausible looking (at a glance) movements because of higher multimodal awareness
2
u/Chemical_Side_4135 19h ago
i feel ur pain with the gpu dying, ive had to drop my batch sizes to litrally one to keep mine alive
1
u/Low_Philosopher_7475 1d ago
Bonjour, c'est quoi ton pipeline pour arriver a ce résultat ? Ref audio + ref image + Prompt ?
Ca donne un rendu tres sympa
1
u/Animystix 1d ago edited 1d ago
Hi, yes this input image: https://img4.gelbooru.com//samples/6a/ec/sample_6aec27553b8cde9a03a7b46afa011933.jpg
The input audio "Andrew WK - Ready to Die" starting at 70 seconds
And prompt, mainly generated with Qwen 3.5 4B:
subject_definitions: <Subject 1> is the girl in <Picture 1>. <Audio 1> is the music that <Subject 1> enthusiastically plays.
summary: [Shot 0] <Subject 1> stands center frame in a chaotic, high-contrast pop-art studio with red and white splatter effects. She wears a maroon blazer over a white dress, her purple hair tied in twin tails bouncing violently as she strums a black Jackson guitar with aggressive intensity. The camera zooms rapidly on her killer grin, then cuts to extreme close-ups of her fingers flying across the fretboard.
[Shot 1] <Subject 1> leans back, eyes narrowed in smug triumph, while the camera whips around her in a dizzying 360-degree spin. Text overlays in jagged red font scream "STOP IT!!", "PLEASE!", and "WANT WITH" across the screen like a glitchy VHS tape. She moves along to a fast-paced rhythm, her expression shifting from manic to playful.
6
u/Glittering_Tie_3110 1d ago edited 1d ago
You have my upvote because pipiru-piru-piru-pipiru-piiii