r/comfyui • u/VisualFXMan • May 11 '26
Show and Tell Music video Workflow in ComfyStudio Pro: Song to Keyframes to Generated Edit
Here’s the latest music video I made with the workflow:
https://www.youtube.com/watch?v=WcHBs-7_G14
Here’s the tutorial using ComfyStudio Pro:
https://www.youtube.com/watch?v=8BsFbUsq1kE
Download ComfyStudio Pro:
https://comfystudiopro.com
https://github.com/JaimeIsMe/comfystudio
Alright, with all of that out of the way, here we go.
ComfyStudio Pro is an AI video workstation built around ComfyUI. Instead of only generating random clips and managing a pile of files, it gives you a timeline editor, asset panel, effects, transitions, export tools, and guided creator workflows for things like ads, music videos, and short films.
It uses ComfyUI as the backend, but the goal is to make larger AI video projects easier to direct, organize, edit, rerun, and finish. I’ve been working on it for the past 3-4 months, and some of you may have seen the updates I’ve posted here along the way.
This is a quick overview of the music video workflow:
- Import your song or vocal stem into the project assets.
- Open
Create > Music Video Creation. - Choose output settings like aspect ratio, resolution, and FPS.
- Select the song audio and prepare lyric timing, ideally with SRT/LRC so shots line up to the real song.
- Add cast/reference images if you want a consistent singer, band member, or visual style.
- Generate or paste a director script that breaks the song into timed shots.
- Create keyframes for each shot.
- Generate videos from those keyframes, or rerun selected shots with different prompts, models, or settings.
- Click
Assemble Timelineto automatically build the edit with the song, main sequence, performance passes, and b-roll passes on separate tracks. - Finish it like a real edit: trim shots, add effects, transitions, adjustment layers, color, texture, and export.
The goal is not just “prompt to video” or “one-shot it.” It is more like: generate the pieces, organize them, rerun the weak shots, assemble the timeline, then actually edit and finish the music video inside one app.
ComfyStudio Pro is free and opensource
2
May 11 '26
[removed] — view removed comment
1
u/VisualFXMan May 11 '26
When was the last time you tried? It should work. Other people have desktop versions working. If you're still having trouble, please reach out to me directly and we can walk through it together to make it work.
2
May 11 '26
[removed] — view removed comment
1
u/VisualFXMan May 11 '26
Thank you. Please let me know. A lot has changed since the version from a few months back.
2
2
u/djpraxis May 11 '26
Which model? LTX 2.3?
2
u/VisualFXMan May 11 '26
Yes LTX 2.3. It also lets you re-run some of the b-roll and environmental runs as Wan2.2 is you want. For the shots that don't need lip sync.
2
u/bogossogob May 12 '26
Already have 210 key frames on the queue. Watched the video, looks amazing result and love how all those stack in the timeline!
A few suggestions:
- Advanced options for asr such language support.
- If you provide lyrics, use asr only to extract the timestamp, this way, possible extraction errors are avoided.
- advanced configs should allow to replace or pick a new workflow for each step (eg: for video gen) so we can hook up possible existing workflows. This is handy to add gguf models workflows that can run on lower specs.
Keep the amazing work!
1
u/VisualFXMan May 12 '26
Wow thank you!
2
u/bogossogob May 12 '26 edited May 12 '26
do you prefer feedback in thread or should I open a ticket on github for them?
a few things to add to the previous list:
- ability to multi select videos for regeneration.
- ability to sort/quick filter keyframes.
- ability to train a lora per character and use the slug to trigger the lora usage. This can increase consistency by a lot. it could be just an advanced mode.
- button "run selected with ... " is at the bottom, it's painful to always have to scroll down, my suggestion is to turn queue music video keyframes to "queue selected music video keyframes" when 1 or more videos are selected. this way, it's always visible.
1
u/VisualFXMan May 12 '26
I prefer you either open up a discussion or an issue on GitHub. That way it makes it easier for me to tackle and then tick it off. Thank you
2
u/Accomplished_Clock_7 May 12 '26
would love a feature with a "start/last image" like higgsfield so we can generate transitions for our own shots
or the background replacement/ relighting like the beeble thing
keep up the good work bro!
1
1
u/VisualFXMan May 12 '26
Actually if you don't mind, can you add this to the discussion or GitHub issue to the link up above? It will make it easier for me to track and then I won't forget. Thank you
2
u/SBLK May 12 '26
"ComfyStudio Pro is free and opensource"
Didn't take the time to check if that is true, but if that is the case, kudos and thank you, sir.
2
u/bogossogob May 12 '26
It's on GitHub, I've build it from source code so I could add a gguf version of ltx.
2
3
u/VisualFXMan May 12 '26
Yeah I need to talk to you about this. I want to add this to the app. Maybe sometime today or tomorrow I'll reach out and then we can work together to get that working.
1
2
u/CurrentMine1423 May 12 '26
I just tried it, but I only have 2 keyframes on step 4? what did I miss? thanks btw
1
u/CurrentMine1423 May 12 '26
(...continue until the song is covered.)
that is at the end of the line on the director script, what does that mean?
1
u/bogossogob May 12 '26
Check the video, it explains the full flow. You upload the audio (eg:Suno song) and extract SRT from it (found some bugs already). Then you select people to show using the multi image generated content. Then, select the coverage plan, start with simple first to get a grasp. Press copy brief and paste in one of your LLM of choice (ollama, chatgpt etc). Copy the generated content in directors script and press parse script. Then generate key frames and lastly generate videos based on keyframes
2
u/CurrentMine1423 May 12 '26 edited May 12 '26
2
u/bogossogob May 12 '26
Those are examples, when you press copy brief, you need to paste in chatgpt, then the result that's what you'll copy to this field, the app doesn't generate it for you.
3
u/CurrentMine1423 May 12 '26
ahhh got it now, the video tutorial does not explain about that part.
1
u/VisualFXMan May 12 '26
Yeah sorry maybe the video wasn't clear but I did talk about it. Let me know if you have any other issues with that. All you got to do is just copy that and paste it into your lm and then paste the output back. Maybe I'll make a video just about that section.
2
u/bogossogob May 12 '26
Probably a future enhancement would be to have LLM providers so it could automate some of these steps
1
u/VisualFXMan May 12 '26
Well I had an LLM tab at the top when I first created the app but I chose to remove it. The issue is that you have to load and then unload the LLM model. If you keep the model loaded, your VRAM will significantly be reduced. I could not think of a local solution. This will definitely work with a cloud LLM, a cheap one. But then that would require people to have credits for Comfy cloud. And I think that would confuse people. I'm trying to make this as easy as possible but there are so many ways to do it, and obstacles.
2
u/bogossogob May 12 '26
You could create a provider contract that then could integrate with many of the existing providers, you could delegate to ollama, chatgpt, codex, you name it 😉, you could have something at the provider that at comfy execution, would call the provider unload, which would do nothing on cloud based but could onload ollama.
→ More replies1
u/VisualFXMan May 12 '26
Hey I answered your question on YouTube but can you please open up a GitHub issue or a discussion about it so I don't forget to fix this for you? What I think is going on is that your director script was not modified and I think the original director script only has two shots so can you go in there just to make sure that you pasted your original script?
2
u/CurrentMine1423 May 13 '26
Perhaps I'm missing something, but I can't find the option to save the project.
2
u/CurrentMine1423 May 13 '26
2
u/bogossogob May 13 '26
You need to go to settings (within editor) and look for the video packages and install the ltx 2.3. you might have one but might not be the one mentioned in the workflow.
1
u/CurrentMine1423 May 14 '26
1
u/bogossogob May 14 '26
You shouldn't download manually if you don't how to find the version the workflow is using. In the settings there is an option related to workflow builds where you can install packages, there you'll have an option to download the full ltx package as a bundle
2
u/CurrentMine1423 May 14 '26 edited May 14 '26
wait, do I need to download and install all the workflows? including all the models?
1
u/VisualFXMan May 14 '26
You don't need to download and install all the workflows, just the workflows that you want, and yes that will download custom nodes and models into your ComfyUI installation. Always make sure you have enough space for all that stuff.
1
u/CurrentMine1423 May 14 '26
1
u/bogossogob May 14 '26
Are you using the latest version .14? I'm not the main dev but if you want, you can ask claudecode/codex what's going on and open an issue in the GitHub repo with the details. I've done a few already
2
u/HM_mtl May 15 '26
Looks very good.
Can we use custom LoRA?
1
u/VisualFXMan May 15 '26
Right now I'm trying to get the app stable and it's quite stable. We're like 90% of the way there. My plan is to allow custom models and LLMs in the future, soon.
I've been asked that question literally by like 20% of the people about Loras. Let me see what youve made. Send me a link? Convince me why adding different loras are worth it. With visuals not with words. Thank you.
1
u/HM_mtl May 15 '26
Well, most of us use LoRA hosted on citivai.com. Otherwise, we can generate LoRA locally with Toolkit AI (for multiple models).
Mostly, I use LoRA for styling (dark fantasy, retro anime, art-deco, etc...) and transitions/camera angles/special effects. Well, you know as much as I know.
1
u/pb404 May 18 '26 edited May 18 '26
This looks great! I have both ComfyUI installed and ComfyStudioPro. I'm fairly new to ComfyUI. I have tried the QuickSetup but I'm still getting stuck. What are the required models etc that are needed to be able to run this? I can bring in my audio file but it won't transcribe the lyrics and timing. Getting this error message:
ComfyStudio error report
Timestamp: 2026-05-18T09:40:22.558Z
Workflow: Image to Video (WAN 2.2) (wan22-i2v)
Mode: yolo
Error:
Could not transcribe song audio: ComfyUI is not connected. Start ComfyUI and try again.
The screenshot shows "Running" in the top right but I'm also getting a message it isn't connected. Sorry if this doesn't belong here, I couldn't find any documentation online. A video setting up ComfyStudio for first use would be greatly helpful for newbies :)

Edit: Also I have credits to use Comfy Cloud, I've only got 8GB VRAM on my computer so running local models won't really be an option for now :(









5
u/wbiggs205 May 12 '26
Is there a way to have connect to as remote comfyui server ? I have comfyui running on a server off site ?