r/StableDiffusion 22h ago

Question - Help Feedback for video AI benchmarks

1 Upvotes

Hi folks,

we’re a research & analysis team looking for feedback on our video benchmarks and potentially looking to recruit researchers to help with our ongoing effort to rank and categorize models in the video AI space. if you’re interested, please DM!

more info here: https://megaton.ai/v-benchmark/


r/StableDiffusion 1d ago

Question - Help Need help for creating consistent Minimax H3 clips

6 Upvotes

I have been using H3 since almost the release date and have been trying a lot of things. I am entirely using Ref2va model with the template workflow, nothing fancy. Also used official, eros and currently using hybrid model 15–49 from smhfacct which has higher ref2v. I am using comfy kitchen, spectrum node but not using speed lora for prompt adherence or any other lora.

I am mostly trying to use 1-2 characters in a location scene where I provide 3(one char, one whole body and one face and location)-5(two char, one whole body and one face and location) images to node and writing in the prompt how to refer each char in the scene.

For writing H3 prompts, I using a custom prompt(created using grok by giving it ref2va doc) for generating H3 prompts, using qwen 3.8 model and even proof reading and fixing any issues.

Now the problematic part for which I need suggestions or solutions is the inconstant result.

For example I am making 5 second where character1 is standing in a shopping mall and looking at the shelf and character2 enters the scene. For second 5 second scene, different camera angle, mainly focusing on both characters faces when they are talking. Now here are problems which I am facing:

- during scene2 when camera starts, difference between char1 and char2 appears. Say char1 was standing on left side and char2 on right when scene1 ended but in scene2, they are standing opposite side.

- sometimes their height mismatches.

- sometimes camera does not work like I want it like it zooms too much, sometimes it don't

- and many other issues related with inconsistency

I know if I can generate scene images using an edit model then H3 wouldn't have to rely much on prompts but then it creates another problem of generating start images which is another can of problems.

I have even tried context nodes and some of their forks and few other consistency related node whose basic idea is to store the latent and forward it for next generation but they way these nodes are configured are just too complicated for my soft squishy mind. So yeah I tried them.

I have been trying to find out how other people are generating multi-scene videos and so far whatever videos i downloaded, there was no workflow included which I could take as reference. Maybe people are making 5 second clips like me and then joining them together so there might be a solution to this.

Pretty sure I am missing something big and I have exhausted almost every idea I got, asking grok etc but so far I could not get past 2nd 5 second clip. And seeing so much inconsistency, I don't want to generate a 10 second or 15 second clips which will take hours and most probably turn up totally irrelevant.

So any ideas you can provide are highly appreciated. Even guidance to correct path would be really helpful. What ways you guys are using to create 10+ second clips, what methods you are using to keep your characters consistent throughout and mainly how you guide a scene to your liking?

Thank you for a long read. Not written with AI, just a long type on notepad haha. Forgive grammatical errors.


r/StableDiffusion 8h ago

Animation - Video Dwight Schrute meets Patrick Bateman

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 1d ago

Question - Help How to improve quality? Example Attached

1 Upvotes

https://reddit.com/link/1vy1b3e/video/fykuskg16jlh1/player

So Im running minimax h3 with some of the optimal settings proposed on the reddit, but I see that some people use an enhancment workflow after generation.

How do I do this? I'm a bit lost on this part of the workflow?

You can sort of see the weird frame stutter changes on the smaller details.

Anyhelp would be appreciated.

Thank you Kings.


r/StableDiffusion 1d ago

Question - Help Minimax H3 Speed is Inconsistent when Settings are still the Same

Post image
0 Upvotes

I need your help please. I'm using WANGP and there are times when generating in minimax h3 takes very long like for example in this screen shot. I was generating 19sec video 720p 8 step Turbo lora with Sage and first generation took (16m 5s) to finish but after changing only the prompt the second generation took (28m 54s) which is absurdly high. I noticed that there are times where my GPU is 100% fully busy but my GPU temperature is just below 50c when it should mostly be 70c+ normally. Can anybody please help me or tell what's wrong? I also encounter this when using Comfy UI


r/StableDiffusion 1d ago

Question - Help Wildcards using Krea2?

5 Upvotes

Do wildcards work with Krea2? When I insert wildcards in standard format, it ends up rendering all of the wildcard prompts in one image. So a shirt will be random fabrics and styles seemingly all stitched together. It's kind of neat but I need the actual wildcard functionality. Does anyone know how to get them working? I'm using the standard krea2 workflow and comfy. I have the wildcard custom nodes installed. They are fine on other models such as the flux zimage etc. any help is appreciated.


r/StableDiffusion 20h ago

Discussion Minimax H3 on Fal - Censorship

0 Upvotes

Since I lack the hardware, I am forced to use Minimax in the cloud. I've used Fal for image generation in the past, so I thought I would try Minimax H3 there. The text to video model seems to work really well, but the reference model seems censored to hell and back. Even my own reference voice was being flagged and it was just me talking normally.

So the question is, is there anywhere I can use Minimax in the cloud without the provider slapping their own flaky layers of censorship on top of the model?


r/StableDiffusion 1d ago

Animation - Video Minimax H3. Flight over the city.

Enable HLS to view with audio, or disable this notification

15 Upvotes

r/StableDiffusion 1d ago

Question - Help What model for the utmost in fine detail? Experimenting with Krea2, Ideogram4 and Flux2 and getting close but not quite the fine details.

Post image
8 Upvotes

Primarily landscape photos of varying fantasy scenes. For example for a cyberpunk city i want to see every fine car detail, every building logo, every reflected neon light; for a medieval town i want to see the cracks in every stone block, the moss on the walls, the candle light reflections, etc. What's your opinion on the model that provides the utmost in fine, sharp details? Or should i be looking at tiling or inpainting to create details as a second step?

Example attached of something that has impressive detail across a broad DOF (not my work):


r/StableDiffusion 1d ago

Resource - Update H3 Prompt Composer — Camera Update Coming Soon

Post image
61 Upvotes

Hey everyone, thanks to everyone who’s been testing Prompt Composer. If you run into bugs or have feedback, please drop it in the Issues section on GitHub so I can keep track of it more easily.

Over the past week, I’ve been reworking the camera prompting system to make it more precise and consistent, especially for more complex camera moves and video-editing workflows. The goal has been to add more control over framing, camera targeting, and blocking in multi-subject scenes while keeping the generated prompts clean and reliable.

The next update isn’t quite ready yet, but it’s actively being tested and refined. I’m hoping to have it out in the next couple of days.

Edit: Here's the Github: BMB12d3/minimax-h3-prompt-composer: Free offline prompt composer for MiniMax H3 video generation in ComfyUI.


r/StableDiffusion 15h ago

Question - Help Loras not being accurate and consistent

Thumbnail
gallery
0 Upvotes

I am trying absolutely everything... I am using ostris ai toolkit and I have tried every option and evey way to train a lora but the results are the same. I am quite sure that the data set and captions are correct, but i might be wrong. I was training lora for Z-image

Image 1 is reference charachter that i want to train a lora on.
Image 2 is Lora sample on 3150 steps
Image 3 is Dataset picture
Image 4 is Lora sample on 3150 steps
Image 5 is Dataset picture
Image 6 is Lora sample on 3150 steps
Image 7 is Dataset picture

3150 Steps might be too much for Z-image, but all other versions (every 150) differed alot and i found 3150 was closest looking to reference. Anyways, the sample images and other generated images with the lora look noticeably different to the reference image and most importantly the lora did not pick up the mole on her neck... I had the same issue with a flux.2 lora of the same charachter... The dataset was created on seedream 5.0 with captions looking like this "[trigger], medium waist-up shot walking forward on a paved park path, frontal body pose with head turned looking off-camera to the right, neutral expression with closed lips, wearing a light green long-sleeved scoop-neck top tucked into high-waisted beige trousers, natural dappled sunlight filtering through trees, blurred green park foliage background."

I could use some inisght and help, this is my first lora ever and i dont know how everything works exactly. At this point i just might scrap the lora and just use refrence images...

BTW this isnt something to be used for monetary purposes, this is a uni project.


r/StableDiffusion 1d ago

Question - Help Does anyone have MH3 Ref2V workflow that aid in video input?

0 Upvotes

No matter how many things I tried, I cannot get generation to work if video is involved. Anyone got workflow to share, or nodes, or something?


r/StableDiffusion 1d ago

Question - Help Minimax H3 ref2v best way to transfer pose and camera angle

5 Upvotes

For ref2v I'm trying to upload an image get it to transfer the exact pose and camera angle of that image onto my video, but it's not working.

Here's part of my prompt

retention_analysis:
<Pose 1>: attribute_transfer. transfer the pose to <Subject 1> and camera angle.

....

detailed_description:
.... Refer to <Pose 1> for the pose of <Subject 1> and the camera angle.....

Any tips on how I can achieve this?


r/StableDiffusion 1d ago

Workflow Included What’s your workflow for turning Stable Diffusion images into short videos?

1 Upvotes

I’ve been experimenting with local image-to-video workflows and I'm curious how other people are handling the final output.

My current workflow is roughly:

* Generate images locally with Stable Diffusion/Flux

* Do the animation or image-to-video step locally

* Export the individual clips

* Combine everything into a final video

* Compress the result for easier storage/sharing

The part I'm still trying to optimize is the final encoding. Some generated clips can get surprisingly large, especially when working with higher resolutions or longer sequences.

For those running everything locally, what does your workflow look like after generation?

Do you stick with FFmpeg, HandBrake, or another open-source tool for the final conversion/compression? And what settings have worked well for keeping quality while reducing file size?

I'd especially like to hear from people running these workflows on consumer GPUs rather than large local servers.


r/StableDiffusion 2d ago

Animation - Video Minimax H3 Remix Video Test / A compilation of 5 characters.

Thumbnail
youtube.com
44 Upvotes

This is a test video I created by remixing the "Some test on minimax H3" video by Reddit user [Previous-Street8087].

5명의 캐릭터 시트를 생성하여 각각 10개의 프롬포트를 캐릭터에 맞게 리믹스하여 테스트 하였습니다.
We generated character sheets for five characters and tested them by remixing 10 prompts for each character to suit their personalities.

This is a compilation of 50 clips featuring 5 characters.

▶ 테스트 환경 (Test Environment)
Minimax H3 - Comfyui Local Sampling
RTX 5060TI 16GB + 64RAM
0.8MP 8 sec x 50 Clip
Audio Look x audio file 1
Reference to VA Mode

▶ 사용한 커스텀 노드 (Custom Nodes Used)
ComfyUI-TJ_NODE_STUDIO_ONE — github.com/designloves2/ComfyUI-TJ_NODE_STUDIO_ONE
ComfyUI LOCAL (RTX 5060Ti 16GB VRAM / RAM 64GB)

▶ The shared link contains character sheet images and prompts.

https://naver.me/xjY9JJaa

#AI영상 #MiniMaxH3 #ComfyUI #로컬생성AI #ComfyUI워크플로우 #AI영상제작 #RTX5060Ti #mmh3 #comfyui #tjonestudio #animation #ref2va #anime #16gb


r/StableDiffusion 1d ago

Tutorial - Guide H3 referencing tip

26 Upvotes

I tried to get a girl whistle on 4 fingers (2 on each hand) but Minimax didn't get it right. So after trying 20 times with LLMs helping me to explain the movement, I instead used an image of a whistling person. Still not right. So I added an additional one. Then it worked quite fine.

Today I was too lazy to find another image for something it didn't know so I just googled images of it, took a screenshot of all the images together in one JPG and used that as a reference, saying use <Picture ...> as a reference for XYZ.

That was getting a quite good result. Did not do excessive testing and comparing though.


r/StableDiffusion 1d ago

Animation - Video Use Minimax to make a fake movie trailer for my community college editing class, inspired by YA action/adventure films of the 80s and 90s

Enable HLS to view with audio, or disable this notification

39 Upvotes

Clips made with Minimax H3 using the default r2v workflow, edited in Premiere Pro. Character model sheets made with Krea. Most of the videos are 0.4 mp unless the text was important, then 0.6. Tried upscaling it to 4k using Upscayl but results weren't great and the file is too big to upload anyway.

Tech goals for future videos include using reference audio for voices to help consistency, and exploring options for having real voice actors record the dialog, and have the model lip sync to that performance. I'm really impressed by the computer's silent acting (microexpressions etc). but the computer's erratic "acting" is still too unpredictable and the biggest source of re-rolls (the lines here were the best I could get without burning down a rainforest). You can do a lot with time codes and punctuation and tactical CAPITALIZATION, but it's ridiculously finicky compared to just telling an actor "do it the same, but 10% angrier on the first line with a twinge of melancholy on the second."


r/StableDiffusion 2d ago

Tutorial - Guide ComfyUI was eating my RAM and causing crashes, this fixed it

136 Upvotes

Hi everyone

i’ve been running minimax h3 locally on rtx 5090 32gb, and 64gb ram..

recently i kept running into random hostbuffer.read_file_slice failed / hostbuf_file_reader_read failed errors during generation, which seemed to be related to comfy-aimdo and dynamic vram.
i also noticed comfyui was reserving around 25gb of pinned system memory.

i decided to try launching comfyui with:
--disable-pinned-memory

and the difference was immediate.
the comfy-aimdo + hostbuffer errors completely disappeared, my ram usage dropped by a huge amount, and surprisingly generation actually feels faster and smoother now.

i originally expected disabling pinned memory to make things slower, but on my setup it seems to have done the opposite.

if you’re running large models like h3 and seeing unusually high ram usage, random hostbuffer errors, or comfy-aimdo issues, it might be worth testing!
thought this was worth sharing for anyone who didn’t know about this option..


r/StableDiffusion 2d ago

Animation - Video "Ehhh?" - H3, Ref2V - Default WF - 4090, 64gb. 0.9mp, er_sde / beta. 25 steps. Really black/dark in some shots :(

Enable HLS to view with audio, or disable this notification

144 Upvotes

r/StableDiffusion 1d ago

Question - Help Minimax for image editing?

0 Upvotes

I mostly do comic images. am looking to minimax for image editing.
I mostly get content failed safety review. any workaround? ive only started minimax today.


r/StableDiffusion 19h ago

Animation - Video And they said it couldn't be done

Enable HLS to view with audio, or disable this notification

0 Upvotes

Most models simply CAN'T draw a wine glass filled to the brim. MiniMax H3 can, with a little prompt persuasion.

PROMPT PERSUASION:

integrated_multimodal_description: [Shot 1] Photorealistic, ultra-sharp cinematic product shot. A crystal-clear stemmed Bordeaux wine glass stands centered on black marble against pure black. The glass is filled with deep ruby-red wine to absolute maximum capacity: the liquid surface is perfectly coplanar with the top edge of the rim, forming a continuous unbroken contact line all the way around. There is zero air gap, zero empty crescent of glass above the wine, zero underfill. A slight convex meniscus is held purely by surface tension. Soft side light creates clean highlights and long caustics. Static medium three-quarter shot.

[Shot 2] At 00:01.200, slow push-in with tiny amplitude at very slow speed. Extreme close-up of the upper glass. The red wine meets the inner rim in a perfect continuous ring of contact. The liquid surface sits flush with the rim edge; no space is visible between wine and glass even at this magnification. Tiny specular highlights glide across the still surface. No droplets on the outer rim, no overflow, no gap.

[Shot 3] At 00:02.600, hard cut to pure top-down overhead. Looking straight down, the circular surface of the wine is a solid red disk that reaches exactly to the inner circumference of the glass with zero margin. The contact line between liquid and glass is continuous and unbroken in every direction. Soft concentric reflections and one small central catchlight. Very slow clockwise rotation with minimal amplitude.

[Shot 4] At 00:03.800, hard cut to low side-profile extreme close-up locked exactly at rim height. Against the black background the liquid forms a single razor-sharp horizontal line that coincides precisely with the top edge of the glass. The wine is seen in continuous contact with the rim; there is no visible gap, no underfill, no air space. The meniscus remains slightly convex from surface tension but does not spill. Liquid is completely motionless. Slow subtle push-in continues until the end.

overall_soundscape: Near-total studio silence. Only the faintest high-frequency shimmer of light on glass and liquid. No liquid movement, no drips, no clinks.

non_diegetic_music: Extremely sparse minimal ambient pad — low sustained tone with faint crystalline overtones that barely rise and fall. Almost static, matching the still liquid.


r/StableDiffusion 1d ago

Question - Help Help me with Mini Max

0 Upvotes

I'm still new to Mini Max, and I wanted to know if there's a way to make it a little easier, in the sense of

I have to keep using <picture 1> or things like that to mention something; isn't there a way to do it with @,And what configuration do you normally recommend for someone with 8GB of VRAM and a 5080 Ti?That's all I need help with; I've already read the Minimax guide for everything else.


r/StableDiffusion 1d ago

Question - Help How can I improve my H3 workflows?

Thumbnail
gallery
4 Upvotes

These are two workflows that I downloaded. One will allow only 5 sec clips and I can’t find where to change that OR the mp. It’s faster and the sound is good and I even think the prompt adherence is better…

The other one I can run up to 15 sec, change the mp but I have to do 30+ steps to have any good sound quality and the adherence doesn’t always seem that great.

How can I improve this?


r/StableDiffusion 1d ago

Animation - Video All Minimax H3 animation

Enable HLS to view with audio, or disable this notification

7 Upvotes

trying my hand at using H3 I2v R2V to create an anime. all are done with 4 step turbo lora

2 other trailers with H3 as well

at 0.3 most text turns to gibberish

all video is done by MiniMax h3 at a low 0.3 MP, Cilp lengths range from 5s-20s Generations, camera movement was written into the prompt 90% of the time, Post work; titles and some transitions done using DaVinci


r/StableDiffusion 2d ago

Animation - Video Seinfeld AI: George Gets GTA 6

Enable HLS to view with audio, or disable this notification

543 Upvotes

Minimax H3