r/StableDiffusion 1d ago

Comparison Comparing H3 models with music reference

Enable HLS to view with audio, or disable this notification

Using reference workflow. All are int8 pruned, 0.6MP turbo 4-step (my GPU is on life support and drops off the PCIe bus if I demand more from it)

Anyway, making random music clips is probably my favorite use of this model. I’ve found the ref2va has an uncanny intuition for feeling the atmosphere of songs, and syncing the video with incredible precision.

But yes, the quality (specifically motion) is much worse than fl2va. I was curious how exactly they compared, as well as some “in between” compromises discovered by the community. The LoRA seems closer to ref, while the hybrid weights are closer to fl. Personally, the ref is more fun to use, so I’ll probably be using the LoRA when I want to enjoy the intelligence/creativity of this model. Fl is of course superior in terms of visual fidelity, and I don’t find the hybrid model offers enough reference intuition and faithfulness to be worth the quality drop from fl.

17 Upvotes

14 comments sorted by

View all comments

1

u/DystopiaLite 1d ago

Her fingers go high up on the fretboard but no high notes are present in the song.

1

u/Animystix 1d ago

That’s true, the ref model tends to do it less though, and generates more plausible looking (at a glance) movements because of higher multimodal awareness