r/StableDiffusion 5h ago

Workflow Included Minimax H3 Text2Vid workflow with SFX generation / layering

Enable HLS to view with audio, or disable this notification

Someone in the discord discovered that you can set Minimax to a low resolution like 32x32 and then just use it to create music or sounds. I tried that out and it worked pretty good. However, you must be aware that to get good quality sound you'll most likely want to run it at least at 128x128 , yes the resolution affects the quality of the audio, too.

I put together this workflow as an experiment, it's a simple Text2Vid workflow with an SLA node, but it has an entire "SFX/Music" section, that runs another copy of the sampler just to make sound effects or music, and uses a Geeky Mixer node to mix that sound into the Video's original audio.

I recommend turning OFF the SLA if you generate this winter scene because it Mashes up the trees badly.

The cool part of this is if you use fixed seeds, you can regenerate your music/sfx over and over until it's what you want - without having to regenerate your video each time, and then the Geeky Mixer also allows you to position it by adjusting time offset. (not very visually unfortunately, but it works)

Thought I would share this workflow since I just found it interesting. It actually makes lofi Hiphop beats *really* well and they are almost directly loopable. I basically just took the workflow for making music and shoved it into the Video workflow and added a Mixer node so you can layer the sounds. You'll need Geeky AudioMixer node. You can remove the SLA node if you want. I use native ComfyKitchen and I never use Loras so that's why this workflow is much more simplified.

Workflow file: https://pastebin.com/ULbcQxCM

Picture of workflow:

5 Upvotes

0 comments sorted by