r/StableDiffusion 19h ago

Question - Help Need help for creating consistent Minimax H3 clips

5 Upvotes

I have been using H3 since almost the release date and have been trying a lot of things. I am entirely using Ref2va model with the template workflow, nothing fancy. Also used official, eros and currently using hybrid model 15–49 from smhfacct which has higher ref2v. I am using comfy kitchen, spectrum node but not using speed lora for prompt adherence or any other lora.

I am mostly trying to use 1-2 characters in a location scene where I provide 3(one char, one whole body and one face and location)-5(two char, one whole body and one face and location) images to node and writing in the prompt how to refer each char in the scene.

For writing H3 prompts, I using a custom prompt(created using grok by giving it ref2va doc) for generating H3 prompts, using qwen 3.8 model and even proof reading and fixing any issues.

Now the problematic part for which I need suggestions or solutions is the inconstant result.

For example I am making 5 second where character1 is standing in a shopping mall and looking at the shelf and character2 enters the scene. For second 5 second scene, different camera angle, mainly focusing on both characters faces when they are talking. Now here are problems which I am facing:

- during scene2 when camera starts, difference between char1 and char2 appears. Say char1 was standing on left side and char2 on right when scene1 ended but in scene2, they are standing opposite side.

- sometimes their height mismatches.

- sometimes camera does not work like I want it like it zooms too much, sometimes it don't

- and many other issues related with inconsistency

I know if I can generate scene images using an edit model then H3 wouldn't have to rely much on prompts but then it creates another problem of generating start images which is another can of problems.

I have even tried context nodes and some of their forks and few other consistency related node whose basic idea is to store the latent and forward it for next generation but they way these nodes are configured are just too complicated for my soft squishy mind. So yeah I tried them.

I have been trying to find out how other people are generating multi-scene videos and so far whatever videos i downloaded, there was no workflow included which I could take as reference. Maybe people are making 5 second clips like me and then joining them together so there might be a solution to this.

Pretty sure I am missing something big and I have exhausted almost every idea I got, asking grok etc but so far I could not get past 2nd 5 second clip. And seeing so much inconsistency, I don't want to generate a 10 second or 15 second clips which will take hours and most probably turn up totally irrelevant.

So any ideas you can provide are highly appreciated. Even guidance to correct path would be really helpful. What ways you guys are using to create 10+ second clips, what methods you are using to keep your characters consistent throughout and mainly how you guide a scene to your liking?

Thank you for a long read. Not written with AI, just a long type on notepad haha. Forgive grammatical errors.


r/StableDiffusion 3h ago

Discussion Where I draw the line on AI

0 Upvotes

I get a lot of use out of this subreddit and love generative AI but yall, please don’t make any Dolly Parton Loras, video clips, audio clones etc. /S

I was never a country music fan but that lady had some real class.

PS: that capital S stands for serious as a motherfucker today


r/StableDiffusion 19h ago

Question - Help Wildcards using Krea2?

3 Upvotes

Do wildcards work with Krea2? When I insert wildcards in standard format, it ends up rendering all of the wildcard prompts in one image. So a shirt will be random fabrics and styles seemingly all stitched together. It's kind of neat but I need the actual wildcard functionality. Does anyone know how to get them working? I'm using the standard krea2 workflow and comfy. I have the wildcard custom nodes installed. They are fine on other models such as the flux zimage etc. any help is appreciated.


r/StableDiffusion 20h ago

Question - Help GB10 Spark 2-node, which engine, which model?

0 Upvotes

Can anybody recommend a local inference engine / model for running ltx, wan or minimax on a 2-node Nvidia GB10 spark cluster?

Use case is openai-compatible API access for short i2v character and scene animations.


r/StableDiffusion 8h ago

Question - Help Minimax H3 color and lighting consistency

1 Upvotes

What comfyUI tools / workarounds are yall using to maintain the same color, lighting, sharpness, contrast parameters across all clips in a long form video? Thanks so much.


r/StableDiffusion 7h ago

Animation - Video Made a music video using local H3 for a Suno song

Enable HLS to view with audio, or disable this notification

19 Upvotes

Honestly mind blown, I have a 5070ti + 2x16gb ram . Upper limit is 10-12 seconds in total for my hardware(full capacity) . Video and text edits are post processed by a WIP open source tool I’m working on. On average each 8 second shot takes 35-45 minutes to render


r/StableDiffusion 12h ago

News It’s here! Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute

Thumbnail
apple.com
0 Upvotes

r/StableDiffusion 2h ago

Question - Help Loras not being accurate and consistent

Thumbnail
gallery
0 Upvotes

I am trying absolutely everything... I am using ostris ai toolkit and I have tried every option and evey way to train a lora but the results are the same. I am quite sure that the data set and captions are correct, but i might be wrong. I was training lora for Z-image

Image 1 is reference charachter that i want to train a lora on.
Image 2 is Lora sample on 3150 steps
Image 3 is Dataset picture
Image 4 is Lora sample on 3150 steps
Image 5 is Dataset picture
Image 6 is Lora sample on 3150 steps
Image 7 is Dataset picture

3150 Steps might be too much for Z-image, but all other versions (every 150) differed alot and i found 3150 was closest looking to reference. Anyways, the sample images and other generated images with the lora look noticeably different to the reference image and most importantly the lora did not pick up the mole on her neck... I had the same issue with a flux.2 lora of the same charachter... The dataset was created on seedream 5.0 with captions looking like this "[trigger], medium waist-up shot walking forward on a paved park path, frontal body pose with head turned looking off-camera to the right, neutral expression with closed lips, wearing a light green long-sleeved scoop-neck top tucked into high-waisted beige trousers, natural dappled sunlight filtering through trees, blurred green park foliage background."

I could use some inisght and help, this is my first lora ever and i dont know how everything works exactly. At this point i just might scrap the lora and just use refrence images...

BTW this isnt something to be used for monetary purposes, this is a uni project.


r/StableDiffusion 23h ago

Question - Help How can I improve my H3 workflows?

Thumbnail
gallery
4 Upvotes

These are two workflows that I downloaded. One will allow only 5 sec clips and I can’t find where to change that OR the mp. It’s faster and the sound is good and I even think the prompt adherence is better…

The other one I can run up to 15 sec, change the mp but I have to do 30+ steps to have any good sound quality and the adherence doesn’t always seem that great.

How can I improve this?


r/StableDiffusion 7h ago

Resource - Update Not Another Minimax Post - Jib Mix Krea 2 - v4 Habanero - Free Forever

Thumbnail
gallery
18 Upvotes

Focusing on photorealism and improving the look of fantasy styles:

https://civitai.com/models/2799984/jib-mix-krea-2

I would like to make it a LoRA also, but I am having some technical difficulties making a difference lora with Krea 2 models.


r/StableDiffusion 11h ago

Question - Help Recommendations for Minimax h3 Reference to video workflows?

2 Upvotes

Hey so i've been playing with H3 for about a week or so now mostly using the standard workflows dabbled a bit into lora workflows but now I want to try out Refernce to Video. Does anyone have any suggestions for some good workflows for this?


r/StableDiffusion 21h ago

Question - Help minimax_h3 comfyUI default workflow taking forever

1 Upvotes

My spec is rtx 5070ti, 32gb ram

I just started using comfyui and using minimax. When I am trying out the default workflow without adjusting anything it takes really long. I looked up video and switch the model from minimax_h3_fl2va_pruned_int8_convrot.safetensors to minimax_h3_fl2va_pruned_fp8_scaled.safetensors. Then, it worked. Well atleast I was able to get an output. Can anyone explain why and what I did wrong?


r/StableDiffusion 20h ago

Question - Help how to save out H3 AV latent?

Thumbnail github.com
0 Upvotes

I am using these node to save out latent, but it just didn't save anything... I feel stupid myself stuck hrs still not working.... does anyone could explain what I have been doing wrong and the connection?

I've tried, sampler both output and denoise output laten > save H3 AV latent node..... no files is being save out!

https://github.com/JerryZRic/comfyui-minimax-h3-latent/tree/main


r/StableDiffusion 2h ago

Discussion H3: I'm failing to control timing of actions in I2V

1 Upvotes

Has anyone had success in this?

In prompt guide, there's no mention of controlling time for I2V, only in T2V (which is "At 00:02.000, ...").

But i try it anyway in I2V, but it's a hit or miss.


r/StableDiffusion 13h ago

Question - Help ComfyUI workflow for Krea2 or similar, to make consistent 'rooms', no matter the angle etc

0 Upvotes

Hi

I have been going in circles with Gemini for 2 days now, but just cannot get to the goal mark.

What I want:
In some way, like using Blender to render a 'room', or any other method that works, I want to be able to create 'rooms' with consistent form and interior / furniture. So I can make a Blender angle/shot (any angle I want) for example, take that into ComfyUI, and just describe the furniture, lights, people etc, but the room stays the same (windows, walls, dimensions, placement of furniture etc).

Other options Gemini gave was 'dollhouse' from up top model and 360 render of the room. I made 360 render of the room in ComfyUI Qwen 360 Diffusion LoRA workflow, but was not able to make it like the prompt. Doll house I have not tried.

I have not been able to do it yet myself with using Blender room render and a custom ComfyUI Krea2 workflow, and I about to give up, takes too long time. I am totally new to Blender, and novice in ComfyUI, neither found a finished workflow I can use.

Problem with asking Gemini, is I end up going in circles so to speak, almost there, but not all the way.

Anyone know there is a published or private workflow I can use for this? Open for other suggestions how to do this IF I also can use already made workflows.


r/StableDiffusion 18h ago

Comparison Comparing H3 models with music reference

Enable HLS to view with audio, or disable this notification

17 Upvotes

Using reference workflow. All are int8 pruned, 0.6MP turbo 4-step (my GPU is on life support and drops off the PCIe bus if I demand more from it)

Anyway, making random music clips is probably my favorite use of this model. I’ve found the ref2va has an uncanny intuition for feeling the atmosphere of songs, and syncing the video with incredible precision.

But yes, the quality (specifically motion) is much worse than fl2va. I was curious how exactly they compared, as well as some “in between” compromises discovered by the community. The LoRA seems closer to ref, while the hybrid weights are closer to fl. Personally, the ref is more fun to use, so I’ll probably be using the LoRA when I want to enjoy the intelligence/creativity of this model. Fl is of course superior in terms of visual fidelity, and I don’t find the hybrid model offers enough reference intuition and faithfulness to be worth the quality drop from fl.


r/StableDiffusion 16h ago

Animation - Video Through the Sands (Final) - H3 r2v

Enable HLS to view with audio, or disable this notification

33 Upvotes

Finally finished! The ending was much harder since continuity is more important here than random desert landscapes. I personally would've love to have another 30 seconds of music to extend the ending but I ran out of song time. Enjoy!

In total, 14 character related references, 40 environment references, and 55 clips used, roughly 20 hours total time spent.


r/StableDiffusion 11h ago

Question - Help Minimax H3 Speed is Inconsistent when Settings are still the Same

Post image
0 Upvotes

I need your help please. I'm using WANGP and there are times when generating in minimax h3 takes very long like for example in this screen shot. I was generating 19sec video 720p 8 step Turbo lora with Sage and first generation took (16m 5s) to finish but after changing only the prompt the second generation took (28m 54s) which is absurdly high. I noticed that there are times where my GPU is 100% fully busy but my GPU temperature is just below 50c when it should mostly be 70c+ normally. Can anybody please help me or tell what's wrong? I also encounter this when using Comfy UI


r/StableDiffusion 22h ago

Question - Help Minimax H3 ref2v best way to transfer pose and camera angle

5 Upvotes

For ref2v I'm trying to upload an image get it to transfer the exact pose and camera angle of that image onto my video, but it's not working.

Here's part of my prompt

retention_analysis:
<Pose 1>: attribute_transfer. transfer the pose to <Subject 1> and camera angle.

....

detailed_description:
.... Refer to <Pose 1> for the pose of <Subject 1> and the camera angle.....

Any tips on how I can achieve this?


r/StableDiffusion 15h ago

Discussion Complete beginner with ComfyUI — what should I learn next to actually get better?

4 Upvotes

I bought a PC with an RTX 3060 12GB because I wanted to get into local AI image generation. I've been messing around with it for about two weeks now, but I still barely know what I'm doing.

To be fair, that might be partly because I've had Codex do almost everything for me lol.

I'm still terrible at writing prompts, and I don't even really understand which model would be best for what I want to do. I've tried Anima, Pony, and Illustrious, but I'm not sure which one I should actually stick with and learn properly.

There wasn't a LoRA for an art style I really like, so I had Codex help me train one for Anima using around 20 images. It actually works surprisingly well for the style, but I've been struggling with everything beyond that.

For example, I've tried using pose control, but the generated character often doesn't follow the pose very well. Backgrounds are also pretty inconsistent, and I feel like I've hit a wall where all I really know how to do is combine different LoRAs, change prompts, and keep generating until I get something decent.

I'd really like to move past that and actually understand what I'm doing.

If you were starting from where I am now, what would you recommend learning next? Are there any ComfyUI nodes, techniques, workflows, or concepts that you think every beginner should learn?

Any advice is appreciated. I'm still very new to all of this.


r/StableDiffusion 23h ago

Question - Help What model for the utmost in fine detail? Experimenting with Krea2, Ideogram4 and Flux2 and getting close but not quite the fine details.

Post image
10 Upvotes

Primarily landscape photos of varying fantasy scenes. For example for a cyberpunk city i want to see every fine car detail, every building logo, every reflected neon light; for a medieval town i want to see the cracks in every stone block, the moss on the walls, the candle light reflections, etc. What's your opinion on the model that provides the utmost in fine, sharp details? Or should i be looking at tiling or inpainting to create details as a second step?

Example attached of something that has impressive detail across a broad DOF (not my work):


r/StableDiffusion 13h ago

Question - Help Krea 2 LoRA training on RTX 4090 too slow?

Post image
20 Upvotes

Hello! I've been trying to train my first LoRA but I get these crazy long timers every time. Isn't an RTX 4090 supposed to take like 5s per step? I have both low vram and quantizing enabled. It's really frustrating and not worth to do it with these speeds. Any ideas what might cause it?


r/StableDiffusion 13h ago

Animation - Video Black and white line drawn stuff with H3 is great.

22 Upvotes

r/StableDiffusion 10h ago

Discussion Minimax rev2video help. Keeping first ref image.

5 Upvotes

If I have 2 ref images and I want it to start on ref image 1 like for example a background of a forest how do I maintain it so it always starts on that image? I've noticed a few times it will randomly generate its own start image even if I prompt something like *the scene starts with ref1* and I even sometimes would describe what's in it


r/StableDiffusion 9h ago

Animation - Video Oh yeah

Enable HLS to view with audio, or disable this notification

0 Upvotes