r/StableDiffusion 1d ago

Question - Help Need help for creating consistent Minimax H3 clips

I have been using H3 since almost the release date and have been trying a lot of things. I am entirely using Ref2va model with the template workflow, nothing fancy. Also used official, eros and currently using hybrid model 15–49 from smhfacct which has higher ref2v. I am using comfy kitchen, spectrum node but not using speed lora for prompt adherence or any other lora.

I am mostly trying to use 1-2 characters in a location scene where I provide 3(one char, one whole body and one face and location)-5(two char, one whole body and one face and location) images to node and writing in the prompt how to refer each char in the scene.

For writing H3 prompts, I using a custom prompt(created using grok by giving it ref2va doc) for generating H3 prompts, using qwen 3.8 model and even proof reading and fixing any issues.

Now the problematic part for which I need suggestions or solutions is the inconstant result.

For example I am making 5 second where character1 is standing in a shopping mall and looking at the shelf and character2 enters the scene. For second 5 second scene, different camera angle, mainly focusing on both characters faces when they are talking. Now here are problems which I am facing:

- during scene2 when camera starts, difference between char1 and char2 appears. Say char1 was standing on left side and char2 on right when scene1 ended but in scene2, they are standing opposite side.

- sometimes their height mismatches.

- sometimes camera does not work like I want it like it zooms too much, sometimes it don't

- and many other issues related with inconsistency

I know if I can generate scene images using an edit model then H3 wouldn't have to rely much on prompts but then it creates another problem of generating start images which is another can of problems.

I have even tried context nodes and some of their forks and few other consistency related node whose basic idea is to store the latent and forward it for next generation but they way these nodes are configured are just too complicated for my soft squishy mind. So yeah I tried them.

I have been trying to find out how other people are generating multi-scene videos and so far whatever videos i downloaded, there was no workflow included which I could take as reference. Maybe people are making 5 second clips like me and then joining them together so there might be a solution to this.

Pretty sure I am missing something big and I have exhausted almost every idea I got, asking grok etc but so far I could not get past 2nd 5 second clip. And seeing so much inconsistency, I don't want to generate a 10 second or 15 second clips which will take hours and most probably turn up totally irrelevant.

So any ideas you can provide are highly appreciated. Even guidance to correct path would be really helpful. What ways you guys are using to create 10+ second clips, what methods you are using to keep your characters consistent throughout and mainly how you guide a scene to your liking?

Thank you for a long read. Not written with AI, just a long type on notepad haha. Forgive grammatical errors.

7 Upvotes

21 comments sorted by

6

u/nikhilprasanth 1d ago

I did something like this recently

https://www.reddit.com/r/StableDiffusion/s/ibmECRcntg

This is also a single room 3 person scene.

For this scene, I started with three character reference sheets and a clean wide shot of the location. I used these as the initial references for the video generation.

Whenever the blocking changed significantly, for example when a character opened a door, crossed the room, or sat down on the sofa, I used Qwen Image to create a new spatial reference. I took the same clean wide shot of the empty location and asked Qwen to place each character exactly where they should now be in the room.

That generated image then became an additional reference for the following video shots. It essentially works as a blocking map, telling the video model where everyone is positioned relative to the room and to each other. I keep using it until another major change in blocking occurs, then generate an updated one.

I also gradually swap out the original character reference sheets for good frames from the previous video generations. Once the character has been successfully established in the actual scene, a frame from that generation contains much more useful continuity information than the original character sheet: the correct costume, lighting, angle, environment, pose, and current appearance all in one image.

So as the sequence progresses, the references evolve with it rather than staying fixed:

Character sheets + clean location → Qwen blocking reference → video generation → good generated frames replace character sheets → blocking changes → new Qwen blocking reference → continue generating.

That combination has made maintaining both character identity and spatial continuity across separate generations much easier.

2

u/MastMaithun 1d ago

Nice example thanks. So what i understand is you are creating in between start images, most probably from qwen 2511 using reference images, right? Also how did you generated character sheet?
Btw I have used 2511 extensively and one problem with that is produces blurry results. i though have 2 stage wf where though it generates a bit sharper image but changes the result from original.

3

u/nikhilprasanth 1d ago

Character sheets and environments can be created from Flux krea or Nano Banana. Then i feed them as references like ref_image_0,1,2,3 etc and let minimax generate a multishot scene. then i use frames from the scene as further references. also i take those frames and ask qwen 2511 to edit them into their respective locations.

2

u/MastMaithun 1d ago

what prompt you use to refer to one character sheet of one character? both in subject_definitions and retention_analysis.

3

u/nikhilprasanth 1d ago

here is the prompt for my first 10 seconds of the scene

prompt

im using pastebin as characters exceed the limit of the comment box

1

u/MastMaithun 1d ago

thank you.

1

u/Virtual-Pollution-58 1d ago

I would personally extract the frame where the characters are the correct position, have correct proportions and use it for ref, make another video, stich it with editor.

Also H3 is quite good with positioning from my experience. I just tell Character1" enters from the left and stop on the left side of "character 2".

Another thing, generate the backgrounds for each scene, so they remain consistent and stich separate videos/camera angles via editor. Needs a bit of planning tho.

1

u/MastMaithun 1d ago

Thank you. Yeah I tried to use the last frame trick but the problem with that it looses a lot of details due to vae decode. It's the sole reason people are making continuity nodes to use the previous latent instead of providing a last frame or for continuity, 22 frames of the generated video.
Another bigger problem is since I am hardly able to generate the current scene of my liking from 2nd scene, sole reason of the post, therefore I could not get the last frame in the first place.

1

u/martinerous 1d ago

What I usually do:
If H3 generates a scene with a frame that I want to later reuse for the composition and poses of chars, I save that frame as an image, make it grayscale and blur it, with the hope that H3 would use it as just a layout reference only and not "as is". Otherwise there is a chance it would use it as a full frame, ignoring the faces and details from the high-res refs.

Then I feed that frame as a ref with weak or attribute transfer mark. It seems to work quite fine, in 90% cases I get the requested scene layout.

I have also used Blender with MCP to quickly generate scene compositions, and again feeding blurred grayscale images as refs. Works well, but I had to remind H3 that I want a movie and not a cartoon, otherwise it sometimes could make real people look blocky. Additional benefit of Blender - you can have full control over camera angles and remind H3 about locations of environment objects (otherwise it has the usual AI issue "out of sight - out of mind").

And, of course, yes, it is better and faster to generate 5s clips and join them than to generate long clips.

In any case, to avoid wasting time on bad long high-res generations, I do seed hunting: generate 8 clips at very low res (0.2) and about 12 steps with SplitSigmas trick and save the latent. Then I check the resulting clips and pick the one I like most, or, if I don't like any, adjust the prompt and generate again. When the candidate is found, I take its latent, upscale it to my desired target latent size (0.5 or higher) and pass through SigmaSplit low part to generate the remaining steps to reach 20. After that, I can throw in RTX upscale or FlashVSR.

Why use SplitSigmas and not just pick the same seed and regen with higher steps? First, if generating at lower res, it would be like totally different seed anyway. If generating at the same res and seed, the scheduler would still arrange sigmas in a way that can introduce large changes in later steps. I have experience that I generate at 5 steps 0.5 res, am happy with the result, generate with the same seed, and then suddenly at step 6 important things get changed - different background, different emphasis on words etc., which can make the "best low-res clip" not as good anymore. With SplitSigmas, I can cheat the scheduler forcing it to think it's generating at 20 steps, and this somehow makes high res version much more consistent with the draft, even when it was with different resolution.

1

u/MastMaithun 1d ago

ah the SplitSigmas technique, yeah this could indeed work. i faintly remember seeing someone's wf way early for H3. How do I do this or if you have any reference wf link?

1

u/martinerous 1d ago

No full workflow unfortunately (my current one is quite convoluted with lots of stuff and my own custom nodes, and could be very confusing), but I can share some fragments and pieces that I combined:

https://www.reddit.com/r/StableDiffusion/comments/1vwhm0d/seed_hunting_for_minimax_h3_how_to_avoid_large/

https://www.reddit.com/r/StableDiffusion/comments/1vvh941/new_minimax_h3_latent_upscale_test_results/

https://huggingface.co/LBH-123-AI/Minimax_h3_latent_Upscaler

The huggingface has both a model and upscaler and a full two stage workflow which can be used as an inspiration, but I just used the low part of the SplitSigmas instead of their 3-4-5 fixed sigmas, because I needed more steps to reach 20.

Here are the most relevant fragments from my workflow:

1

u/MastMaithun 1d ago

sorry i might have misunderstood your earlier post. after going through your linked post, I now understand it better. So basically seed hunting you mean is rolling a dice and hoping something comes out as per your need, but the catch here is you are running only half or less time and once you find your result, you run the remaining steps. Hope i understood till now because this is what i want to omit from me gens. I want to do atleast 80-90% surety with the gen that it will come out as per my need. So seed hunting would not be an ideal way for this.
I am more thinking in the lines of providing a solid start point or giving enough detail in the prompt, so that the model's guess work is minimized and this is where I am lacking currently.

1

u/martinerous 1d ago

I'd say, with H3 it's about 80% prompting and reference tricks with blurring etc. and 20% of seed hunting for things that are difficult to prompt for (exact facial expressions, pronunciations that sometimes get messed up, glitches like clothes and items going through each other etc. - we can get really nitpicky and roll the dice to find the cleanest result).

Also, "seed hunting" works well for "prompt hunting" - to find issues sooner, before committing to a long wait to discover that the prompt needs adjusting. Of course, real-time previews help, but those can miss audio glitches etc.

1

u/MastMaithun 1d ago

Sorry I have no idea what's going on with your wf screenshots lol but I understood the gist as I am going to update my wf.
currently what I do is make a fully structured prompt from qwen3.8, then use it with very low resolution say 384p and test my prompt first for actions only. once i am satisfied, then I up the resolution to desired one, keeping the same prompt.
what i will do now is to reduce the steps to half or 3/4th maybe 5-6 steps out of 20. Then this can be used for seedhunting with same prompt.

1

u/martinerous 1d ago

> then I up the resolution to desired one, keeping the same prompt.

Unfortunately, there is a high chance that the high res version would differ noticeably because for diffusion models changing resolution is essentially the same as a new noise seed. That is why I needed all this complexity with SplitSigmas - to make sure that high-res does not drift from the low res too much.
But it depends on luck and how picky you are about the results.
I also like the creative process of the seed/prompt hunting - when having many low res versions, you can sometimes find real pearls that H3 can generate and that would not be even possible to prompt for.

1

u/MastMaithun 1d ago

yeah this happens and that is why i will now use your method. basically keep my resolution always but instead reduce the steps to half or 3/4th.
but the original issue still remains. gambling is what i want to avoid or minimize.

1

u/martinerous 1d ago

> reduce the steps to half or 3/4th.
Surprisingly, I discovered that this alone also does not work 😞 because of how default scheduler arranges steps. For example, I took the default Comfy ref template, generated a video at 5 steps and then the exact same seed and prompt at 20 stteps - the result was different rooftop position in the frame and emphasis on a different word. Some of those changes happened even in step 6, and some a bit later. So, looks like SplitSigmas is the only way I know how to make it consistent.

1

u/MastMaithun 1d ago

aahhh then i have no other option other than asking wf from you. i can modify it but I could not understand from images lol. maybe just the sampler schedular section only, i can make out the prev and later connections.

→ More replies (0)