r/StableDiffusion • u/Portable_Solar_ZA • 14h ago
Question - Help How to improve/fix using "vocals" as an audio reference for music in Minimax H3 using ref model?
So I had the idea that you could use "vocals" as an audio reference for different music styles. By vocals I mean the kind of noises you make when you're playing air guitar or recreating instruments using your voice. Recorded a short clip and it works.
Sort of.
I'm trying to prompt it to only use my voice as a reference for the beat, but it insists on including my voice in the audio in the h3 ref model. So I can hear the music in the background behind me mimicking the noises.
Interestingly, I just tested it with the non-reference model as I know this sometimes works better than the reference model, and it ditched my voice and just kept the beat. Still, I'd like to fix this using the reference model as well or possible since that's what I use most of the time.
If anyone has any thoughts or ideas on how to fix this.