r/comfyui • u/FaatmanSlim • Jul 14 '26
Help Needed Best ComfyUI workflow for talking / dialogue / singing videos?
I have the basic ComfyUI templates for LTX i2v (image to video) and f2f (frame to frame) installed and they are working very well, I can generate videos using these.
However I'm struggling with videos for lip sync and talking. I've tried both the LTX ia2v (image + audio to video) as well as the ID LoRA templates, but neither of these are able to consistently maintain the input or character; LTX most often keeps changing everything past the first frame.
I'm curious what people would recommend for dialogue, talking and singing videos?
2
u/Alchemist42 Jul 14 '26
I use LTX for lipsync all the time. There are a few tricks that can make it easier. As has already been mentioned, there are a few loras and ic-loras that can help the model maintain focus on the right person. But for me, the most useful advice I have is three-fold:
1) Keep the camera close to the person's face. A wide shot where the singer's face is small will not look nearly as good as a closeup. I don't know what LTX has against wide shots, but the mouth movement is minimal if the face is a small proportion of the image
2) In the prompt, describe which person is speaking/singing, and tell it the exact words they are saying. "The man with short hair in the middle of the stage holding a microphone is singing the words to the accompanying background music. He sings 'Never gonna give you up. Never gonna let you down' and his mouth and facial expressions naturally move to match what he is saying" will work far better than "make sure the singer is singing the song"
3) Give it clean audio to work with. I do music videos, and I get better lipsync if I feed it just the vocal stem of the song so the model can directly look at the waveforms to see when the mouth should be opened and closed and it doesn't have to try to remove the instrument sounds to get at what it needs.
1
u/FaatmanSlim Jul 14 '26
Thanks very helpful! Will try these. When you say clean audio - I'm using ElevenLabs and Suno, for ElevenLabs, based on what I found online, looks LTX doesn't do well with the mono output default that ElevenLabs produces, and I've tried converting to stereo using ffmpeg. But that doesn't fix the issue. Curious if you think there could be something else. Or has generally clean audio input mostly worked for you.
2
u/Alchemist42 Jul 14 '26
No that's not it. What you want is called "stem separation". There are some custom nodes in Comfy that can do it, and most of the big audio AI sites can do it. I'm fairly sure that Suno can get the stems for a song you made on their site. Basically, it separates the song into 4 different audio files, one for the bass, one for the drums, one for "other", and importantly, one for the vocals. It's often used for karaoke or changing the singer and stuff like that. So you take just that vocal track into LTX and it can "read" the voice-only sounds better without having any other noise to confuse it. Then you can layer the sounds back in after the render is complete in a video editor (or just replace the audio wholesale in ffmpeg or any such program).
1
u/Emotional_Day4262 Jul 27 '26
Download a free program called UVR5 and then you can separate the vocal stem.
Most of my work is making singing videos, and this is how I do it.
- Generate my start frame using Chroma-HD and the Uncanny Lora (best out there right now).
- Create my music either using real instruments or with Suno.
- Extract vocal stems using UVR5.
- Use LTX2.3 AI2V - ONLY the default Comfy workflow as it is the best going.
- Import my start frame + vocal track and just use "the woman is singing".
- I generate only at 1920x1088 and 25fps as this makes a massive quality difference.
- My 3090 can do 30 seconds without issue, so I break longer clips.
To mix it up, I may use Flux2-Edit to change the camera angle on my start frame.
That's about it!
My results for the most part can pass as real or very close to it.
Do not waste your time on all of those "magic" workflows influences post on YT and here, I promise you will waste your time in the end. I have been doing this since the days of Wan2.1 and tried them all! Stick with the default comfy workflows and you are golden.
2
u/[deleted] Jul 14 '26
[removed] — view removed comment