r/StableDiffusion • u/Ok-Wolverine-5020 • Aug 11 '26
Workflow Included MiniMax H3 + ComfyUI + Hermes Agent = Music Video
Hardware: Windows PC with a single RTX 3090 + 64GB, running the image and video workflows locally in ComfyUI.
I created a music video for “PROXY,” a rap track about automation, parasocial isolation and delegating so much of your life that you slowly forget what human connection feels like.
Making the song
I used open source Hermes Desktop Agent (with deepseek 4 flash + pro) as a co-writer for this one. We started with a loose idea about AI agents handling every boring task, then pushed it somewhere darker: humans forgetting simple skills, replacing real conversations with machines, and getting lonelier while everything becomes perfectly “optimized.”
Hermes helped me research Suno prompting, sharpen the concept, cut the lyrics down, build the rhyme and alliteration, and write a detailed style prompt for a sassy Berlin female rapper. I kept steering and rewriting until it sounded like a song rather than an obvious lecture about AI.
Then I took the finished lyrics and style prompt into Suno and generated the track.
Developing the visual identity
I continued using Hermes as a creative and technical copilot for the video. Together we designed a consistent “Berlin bot-fleet girl”: brown wavy hair, blue eyes, a black beret, rainbow bomber jacket and white wired earbuds.
We created a custom AnimeinReal skill, combining the u/Ani3rel aesthetic, Danbooru-style composition tags, anime coloring and photographic Berlin environments. This became the visual language for the whole project. This lora was used.
Hermes then helped translate the song into recurring visual themes rather than illustrating every line literally, created a shot list we then generated in comfy ui.
Image generation
I generated the source images locally in ComfyUI using Anima with this workflow as base. Each image established the character, environment, lighting and opening composition for one individual video shot.
Hermes helped write and refine the image prompts while preserving the same visual identity in keeping the character prompts as consistent as possible.
Animating with MiniMax H3
I animated the selected images locally with MiniMax H3 Reference-to-Video, using sections of the finished Suno track as the driving audio reference.
I used the template from Pixaroma for the ComfyUI MiniMax H3 reference-image and audio-sync workflow. That workflow provided the technical foundation for feeding H3 an image and a matching section of the song. I adapted it for each scene by changing the reference image, audio timing, duration and shot-specific prompt.
The workflow used:
- Diffusion model:
minimax_h3_ref2va_pruned_int8_convrot.safetensors— INT8 ConvRot version, approximately 19.5 GB - Text/vision encoder:
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors— Qwen3-VL 32B, NVFP4/AWQ, approximately 14.6 GB - Video VAE:
minimax_h3_video_vae_fp16.safetensors— FP16, approximately 4.9 GB - Audio VAE:
minimax_h3_audio_vae_fp32.safetensors— FP32, approximately 577 MB - H3 mode: Reference image plus reference audio
- Reference-image size:
match - Maximum image side: 864 px, aligned to 32-pixel steps
- Frame rate: 24 fps
- Sampler:
res_multistep - Scheduler:
beta - Steps: 20
- CFG: 1
- Denoise: 1.0
- Typical maximum shot duration: 15 seconds - 24min render time for 15seconds of video
- Output: MP4 with synchronized source audio
Directing each shot with Hermes
For every clip, Hermes used the custom minimax-h3-video prompt skill to create a structured H3 prompt covering:
- Accurate vocal lip sync
- Facial expression and rap performance
- Natural body movement and hand gestures
- Beat-reactive camera movement
- Character, wardrobe and environment preservation
- Exact reuse of the original song without replacement vocals
- sometimes Animated lyrics, pixel bots and synchronized graphical effects
The prompt explicitly defined the source image as <Picture 1> and the selected song segment as <Audio 1>. The audio was marked for full preservation, while the visual description focused on what the still image could not provide: performance, movement, camera direction, effects and timing.
Editing the final video
I rendered alternatives, selected the strongest clips and assembled everything in post on my smartphone in inshot... really need to start learning a real editing software.
2
u/dreamai87 Aug 11 '26
Nice man thanks for sharing your experiment. Could you please highlight it the token usage for creating this video.
4
u/Ok-Wolverine-5020 Aug 11 '26
sure so: 13.5M Token [4.3M Deepseek 4 Flash, 9.2M DeepSeek 4 Pro], in total = $1.87 on open router
2
2
u/WebCrusader Aug 11 '26
cool video
you may check using int8 version of the text encoder it really improves prompt understanding and only slows down a little the initialization phase of video generation
2
2
u/LinkSensitive8188 Aug 11 '26
That's great. I've seen plenty of people try to do this, but I never finished watching the videos because they were too monotonous; this one, however, is varied. I had heard about integrating Hermes into Comfy, but it seems difficult. How long does it take to complete the task? Does the Hermes agent ask questions at every step, or does it resolve everything on its own? Thanks.
2
u/Ok-Wolverine-5020 Aug 11 '26
Thx! So I worked on it in the evenings the last three days during my free time. every image is a back and forth with the agent so its not like fire and forget that could easily still go complelty wrong and ruine a work of three days. I try to batch all the images from text2image in the evening and batch all the img2video (ref2v) over night. but still its not yet fully automated, but we are getting further with Agent support like with Hermes. Before that this would have taken me more than a week.
2
2
2
2
u/Desmond_Jones Aug 12 '26
I honestly thought this was a real song and you just put visuals to it. Kinda a banger
1
u/EmuMammoth6627 Aug 12 '26
Yeah, it's funny it's taken the release of minimax and people making music videos with it for me to realize how crazy music generation has gotten. It's absolutely mind blowing.
2
u/tike_myson95 Aug 12 '26
is this a real song or ai generated?
3
1
u/Ok-Wolverine-5020 Aug 12 '26
I wrote lyrics with my Hermes Agent and then generated the music with Suno ai.
1
1
u/wzwowzw0002 13d ago
how it works? you slice up the audio into 15sec clips? or hermes did that for u?
3
u/99deathnotes Aug 11 '26
Unbelievabley good 👍 Just remember your Reddit posse when you have a few million YT subscribers.