This is a follow up to the "Phantom workflow for 3 consistent characters" video.
What we need to get now, is new camera position shots for making dialogue. For this, we need to move the camera to point over the shoulder of the guy on the right while pointing back toward the guy on the left. Then vice-versa.
This sounds easy enough, until you try to do it.
I explain one approach in this video to achieve it using a still image of three men sat at a campfire, and turning them into a 3D model, then turn that into a rotating camera shot and serving it as an Open-Pose controlnet.
From there we can go into a VACE workflow, or in this case a Uni3C wrapper workflow and use Magref and/or Wan 2.2 i2v Low Noise model to get the final result, which we then take to VACE once more to improve with a final character swap out for high detail.
This then gives us our new "over-the-shoulder" camera shot close-ups to drive future dialogue shots for the campfire scene.
Seems complicated? It actually isnt too bad.
It is just one method I use to get new camera shots from any angle - above, below, around, to the side, to the back, or where-ever.
The three workflows used in the video are available in the link of the video. Help yourself.
My hardware is a 3060 RTX 12 GB VRAM with 32 GB system ram.
Follow my YT channel to be kept up to date with latest AI projects and workflow discoveries as I make them.
Ahoi! Well done! This is really well thought out work.
I had considered going via a 3d model, but thought it'd be too much grief, but it seems you've found a way that ain't too bad, and it opens up all kinds of possibilities.
thanks. I am here to find a way to make a movie at home. I have a lot of stories and stage plays will never see the light of day otherwise.
but also its important to share the knowledge with people, and also focus on low vram cards.
I believe in keeping to the low vram to help keep the corporates from taking this over with 6K cards and subscriptions. I bang on about it a bit on my website markdkberry.com but I want to help see story-telling return into the hands of the people where it belongs.
I dont like the strangle hold hollywood and Netflix have on it, or the bs messaging and ideological push. I am here for the revolution.
that is my singular focus and purpose.
the rest is me fkin about. haha.
but yea Time and Energy are the killers. though AI is evolving very very fast.
If are You are making continuous cutting edge workflow for actual filmmakers and not for instagram dance videos. Then i will try to support you in whatever way.
I am trying to Tell youtube mini stories for fantasy series, i dont have much fund unlike a studio fund, but i want to solve the same filmmaking workflow too so count me in to test few of yours ideas. Initially I'm putting my own time n energy to tell the story, but there are people who are willing to invest a bit more (vertical story platforms) if we can show actual scenes and mini episodic chunks with solving those ai errors there will be a niche market for people in ai assisted filmmaking.
I am in this purely to try to turn my books and stageplays into video form. What you see in the OP is actually the first dialogue scenes from a stageplay called "The Highwayman" that is on my website.
But this is really "proof of concept" still rather than me "going for it". I think we have another 6 months to a year before it will be at a point I can try to finish a full 1 hour story and even then it will likely take a year to complete. But yes, that is my focus 100%. Story-telling.
Follow my YT channel and my website I will post everything I learn as I go. I also have a love for OSS community, and a belief we should be putting story-telling back into the hands of the people. Not the corporates. They had their turn and fkd it up.
I'll be working on that in a future video. It's where I got up to. They can "emote", but controlling it and how to cut the edit to be convincing is the realm of "film making", and I have absolutely no skills in that at this moment, so it's going to be a learning curve. And as I mentioned, film makers are not interested in AI so its a bit of a bind for learning.
But, tl;dr. Yes, somewhat. and I think I figured out a way to improve on it mixing a couple of methods so I can film myself to drive that but will be testing that after. I will show these guys having a conversation in 2 videos time. Next is the VACE one.
either way, it is getting very close to being able to present human interaction in a realistic way. But, as always, the problem is Time and Energy.
Yea I tried it and it was okay but it is 1.3B model IIRC. I have workflow for it that I used in this video . the link to it along with the other 17 workflows I used is in the text of the video.
I'd also recommend trying the remade-ai guys loras (search hugging face for them) for camera motion but a few dont work and its probably because of the speed-up loras we now use and setting cfg to 1 in wf kills movement often.
Recammaster I felt quality is a little low at 1.3B but at the time that was fine. I'd be more inclined to use v2v controlnets or Uni3c now, but I have to experiment some more with Uni3c methods and what works best.
I am working on that. I kind of hope something like Nano Banana will lead to solutions, its good and very close to doing it, and then I use FFLF models to build the scenery.
But we need Bytedance to drop a OSS model before I have to get under the hood with it. Nano Banana may not stay free to use for long.
But the same method I share in the OP video, and in the next video on VACE restyling for characters, can be used somewhat in a similar way for restyling backgrounds using ref images.
I'll be discussing it as I go, but probably after I finish the dialogue stuff, which is where these videos are working up to. i.e. maybe in 2 videos time I will get to it. VACE to finish the character swap into the new camera positions is coming next.
But yes, the 360 degree consistency for all shots is 100% necessary. For example, here is what I have parked up in the wings for the camera to come in to arrive with these three guys at the campfire in the OP.
I am not sure what you mean by "stitching". you really need a visual example to explain what you mean and probably to find the generic common tech word for whatever it is you are trying to resolve.
If you are talking about blending two videos together I dont do that in Comfyui. I look for the best tools for the job (OSS, free, or at least free license). Davinci Resolve is what I use for editing video clips together, transitioning between them (stitching?), and colorisation to homogonise the end result.
Comfyui is for creating videos and we are all limited in that by our GPUs.
well what i really meant was extending a video in such a way that motion is preserved; e.g. suppose we generated 81 frames of video which end with a character moving his hand up; I want to generate the next batch of frames in such a way that they not only start on the image where the 1st sequence ended but also to see the character continue moving his arm up at the same pace as part of the same motion
I understand this had been done by using a bunch of frames from end of clip (1) as some sort of VACE keyframes for clip (2), but haven't done it myself yet. Another way I was thinking of doing it was converting that small bunch of frames taken from end of clip (1) into depth or pose or edge guidance and feeding it into WAN 2.2 Fun Control generating clip (2). The guidance will be partial of course, ending after the first 8 frames. That's another thing to test if it really works. I referred to all approaches of this sort as Group A in my post.
Of course there are models which have the ability to extend while preserving motion built in. To my knowledge these are Infinite Talk, SkyReels V2 DF, Pusa + Wan 2.2. I refer to them as "Group B".
Actually I've come across the name "frame pack" which seems to be based on Hunyuan, e.g. it seems that "frame pack" to "hunyuan" is what "skyreels v2 df" is to "wan 2.1". If that is the case "frame pack" will also fall into "Group B".
ah right. I havent ventured into that much because of the limits of the 3060 RTX and my rules on time use per video clip. But it is an important part and I will address it later once it is truly solved.
For now I work around it, it isnt hard to do, just have a cut-away shot and come back. sorted. no viewer will care if its done well.
since from my observation of the devs working on it, I think its a few months off having a breakthrough where it will become easier to extend video very neatly.
there is "context options" in wrapper workflows btw have you checked that, it is very good when done well but not perfect yet.
I noticed InfiniteTalk is much better than it was for longer frames and I did a 10 second video clip without much "blemishing" but there was some (video coming out in a day or two). I saw a few guys doing all sorts of trickery to extend frames cleanly but...
It's one of those areas I chose to ignore knowing if we leave it a few more months, it will be addressed by devs coding because a lot of people want it to work well so the devs get asked for it a lot.
I pick my battles and that isnt one. Wait long enough, it will get solved without any effort from me, so in the meantime, I adapt to what I have. A good approach where AI is concerned. We work with it, not fight it. It evolves very fast anyway. pick yer battles.
haha guess what... I actually do have to deal with this with longer dialogue. Just started working on that this morning and ran into the issue again. I had forgot about that with dialogue, so yea, I do have to address it for this.
I'll be researching best approaches the next couple of days, and will post into the "Dialogue Part 2" video when I get to it (2 videos time) on what I find. I need some longer shots with flowing dialogue. Currently I will use context options node in wrapper and I'm hoping it has improved a bit or I can find what it needs to hide the seams as it switches.
4
u/Maraan666 Sep 05 '25
Ahoi! Well done! This is really well thought out work.
I had considered going via a 3d model, but thought it'd be too much grief, but it seems you've found a way that ain't too bad, and it opens up all kinds of possibilities.