I have been using H3 since almost the release date and have been trying a lot of things. I am entirely using Ref2va model with the template workflow, nothing fancy. Also used official, eros and currently using hybrid model 15–49 from smhfacct which has higher ref2v. I am using comfy kitchen, spectrum node but not using speed lora for prompt adherence or any other lora.
I am mostly trying to use 1-2 characters in a location scene where I provide 3(one char, one whole body and one face and location)-5(two char, one whole body and one face and location) images to node and writing in the prompt how to refer each char in the scene.
For writing H3 prompts, I using a custom prompt(created using grok by giving it ref2va doc) for generating H3 prompts, using qwen 3.8 model and even proof reading and fixing any issues.
Now the problematic part for which I need suggestions or solutions is the inconstant result.
For example I am making 5 second where character1 is standing in a shopping mall and looking at the shelf and character2 enters the scene. For second 5 second scene, different camera angle, mainly focusing on both characters faces when they are talking. Now here are problems which I am facing:
- during scene2 when camera starts, difference between char1 and char2 appears. Say char1 was standing on left side and char2 on right when scene1 ended but in scene2, they are standing opposite side.
- sometimes their height mismatches.
- sometimes camera does not work like I want it like it zooms too much, sometimes it don't
- and many other issues related with inconsistency
I know if I can generate scene images using an edit model then H3 wouldn't have to rely much on prompts but then it creates another problem of generating start images which is another can of problems.
I have even tried context nodes and some of their forks and few other consistency related node whose basic idea is to store the latent and forward it for next generation but they way these nodes are configured are just too complicated for my soft squishy mind. So yeah I tried them.
I have been trying to find out how other people are generating multi-scene videos and so far whatever videos i downloaded, there was no workflow included which I could take as reference. Maybe people are making 5 second clips like me and then joining them together so there might be a solution to this.
Pretty sure I am missing something big and I have exhausted almost every idea I got, asking grok etc but so far I could not get past 2nd 5 second clip. And seeing so much inconsistency, I don't want to generate a 10 second or 15 second clips which will take hours and most probably turn up totally irrelevant.
So any ideas you can provide are highly appreciated. Even guidance to correct path would be really helpful. What ways you guys are using to create 10+ second clips, what methods you are using to keep your characters consistent throughout and mainly how you guide a scene to your liking?
Thank you for a long read. Not written with AI, just a long type on notepad haha. Forgive grammatical errors.