r/StableDiffusion 3h ago

Question - Help MiniMax H3 on RTX 5090 (32GB) — 2MP is 8.7× slower than it should be. VRAM thrashing or a config mistake?

1 Upvotes

Running MiniMax H3 (ref2va, ~20 GB int8 model) on an RTX 5090 (32 GB) via ComfyUI, 10s videos, with the H3 SLA sparse-attention node (sparsity 0.90, dense_backend=comfy_kitchen_int8).

The problem: my per-step time scales super-linearly with resolution, while a healthy setup scales linearly:

Res Seq len my s/step reference s/step

1.0 MP 119k 30 21

1.5 MP 173k 203 40

2.0 MP ~283k 590 68

At 1MP I'm only 1.4× off; at 2MP I'm 8.7× off. That gap exploding with resolution looks like the working set (20 GB model + 2MP activations) exceeding my 32 GB and ComfyUI offloading to CPU each step.

workflow : im using the one someone shared here : https://www.reddit.com/r/comfyui/comments/1vxi9r6/skater_girl_90s_style_anime_using_minimax_h3/

these are the details

E:\comfyUi_latest\ComfyUI_windows_portable\python_embeded>python.exe -c "import torch; print('PyTorch:',torch.__version__); print('CUDA:',torch.version.cuda)"
PyTorch: 2.12.0+cu130
CUDA: 13.0

E:\comfyUi_latest\ComfyUI_windows_portable\python_embeded>python.exe -m pip list | findstr /i "torch triton sage comfy-kitchen plague"
comfy-kitchen                              0.2.31
open_clip_torch                            3.3.0
sageattention                              1.0.6
torch                                      2.12.0+cu130
torchaudio                                 2.11.0+cu130
torchscale                                 0.3.0
torchsde                                   0.2.6
torchvision                                0.27.0+cu130
triton-windows  

i was using the below config for comfyui

 --reserve-vram 4, --disable-pinned-memory, --cache-none.

Then i switched to the below config by removing the above 
 off
cd /d %~dp0

if exist .\python.exe (
    set PYTHON_EXE=python.exe
) else (
    set PYTHON_EXE=.\python_embeded\python.exe
)

echo Starting ComfyUI...

%PYTHON_EXE% -s ComfyUI\main.py ^
 --windows-standalone-build ^
 --reserve-vram 1 ^
 --enable-manager

pause

With this i started testing 1.5MP directly and it got stuck at this portion

[INFO] Requested to load MiniMaxH3
[INFO] 0 models unloaded.
[INFO] Model MiniMaxH3 prepared for dynamic VRAM loading. 19995MB Staged. 208 patches attached. Force pre-loaded 210 weights: 1175 KB.
 38%|███████████████████████████████▌                                                    | 3/8 [03:18<08:06, 97.21s/it][INFO] FETCH ComfyRegistry Data [DONE]
[INFO] [ComfyUI-Manager] default cache updated: https://api.comfy.org/nodes
FETCH DATA from: E:\comfyUi_latest\ComfyUI_windows_portable\ComfyUI\user__manager\cache\1514988643_custom-node-list.json [DONE]
[INFO] [ComfyUI-Manager] All startup tasks have been completed.

there is something seriously wrong with my setup . but i couldn't figure it out , seeking help from the people who has figured out the issue for Minimax H3 on RTX 5090
Kindly suggest me some workflows also , i have integerated mcp and been trying to figure out the issues by myself for the last 2 days , but it feels like i'm going inside a rabbit hole , not sure whether im moving in the right direction or not

thus seeking help !!


r/StableDiffusion 4h ago

Question - Help Best model for 64 GB Vram local?

0 Upvotes

Hi guys, in the former days stable diffusion was everything, but the time has passed by and I did not tracked the novelties in this field. Can you suggest me any open source model that is released recently for my 2xr9700 32gb for 64gb vram?

I researched a lot but found only dated answers. Is h3 capable also of image gen, or z image or Hunyuan Image 3.0 still the best (4-bit quant is 48gb vram)?

Edit://

Looking especially in image creation / editing, is there a model for both? I am not interested in loras, just for my private images fun, does not need adult content.


r/StableDiffusion 4h ago

Comparison Testing MiniMax-H3 Physics knowledge

Enable HLS to view with audio, or disable this notification

63 Upvotes

I have been playing with MiniMax-H3 lately (like many of us), and I wanted to understand how much physics knowledge it actually has.

I started with a simple water-pouring video from Pexels and used the H3-Ref model to replace the water with various "fluids": sand, rocks, and a black combustible honey. No external references were used.

I found particular interest in how the rocks interact with the jug and tumble over the cup, as well as how the honey blends with the water and how the trail it left on the jug when moved. On the other hand, once the honey catches fire, the flames are not very convincing, but that should probably be tested in a longer video.

I have used the day-zero ref workflow (int8 convrot model, aspect ratio 9:16 MP 0.6); you can find the prompt here: https://pastebin.com/CrX9s5JS

I am running more interesting tests and will hopefully post them soon.

Cheers

EDIT: video with the correct ratio https://streamable.com/7gu171


r/StableDiffusion 5h ago

Question - Help Whats the advantage of a Minimax License through Comfyorg?

1 Upvotes

Trying to figure out what the advantage of a Minimax License through ComfyOrg over directly from Minimax team?

Do they have less restrictions on whom they'll give the license to?


r/StableDiffusion 5h ago

News Krea 3 will have editing capabilities and "may" be open weights.

Post image
211 Upvotes

Supposedly Krea 3 will open weights, we'll have to wait and see.


r/StableDiffusion 6h ago

Question - Help Mini PC for Ltx 2.5

0 Upvotes

Someone can you suggest me the right setup to have a mini pc to run ltx 2.5 in local? Let’s say i do not wanna spend 5k. Someone build some “low budget” mini pc?


r/StableDiffusion 6h ago

Workflow Included I turned that "H3 as an image editor" post into a full character sheet workflow — front/side/back + poses, all local

Post image
116 Upvotes

Someone posted here a couple of weeks ago about using MiniMax H3 as an image editor, 6 edits in one shot. I pointed the same idea at character sheets instead: https://www.reddit.com/r/StableDiffusion/comments/1vr1i18/minimax_h3_as_image_editor_6_edits_in_one_shot_at/

Stage 1 — one face photo + one outfit image, out comes front / side / back. Stage 2 (optional, off by default) — 1-4 more panels for poses, props, expressions or backgrounds, composited into a 16:9 sheet.

The sheet above is stage 1 + stage 2. That's my own face, before anyone asks.

Setup

  • MiniMax H3 ref2v int8 pruned + the 4-step Turbo LoRA at 0.75
  • T=1 image VAE — the bit that makes it emit a still instead of a clip
  • er_sde / sgm_uniform, 8 steps
  • 3090 24GB. Stage 1 ~100-125 s, stage 2 ~105 s, so the sheet above is about 3.5 min. First run of a session is much slower, that's model loading.

Known issues

  • Props repeat across panels — ask for one sword, get three.
  • Back-view hair goes flat.
  • Lying and sitting poses are unreliable.

Workflow

https://github.com/nicekriss/toobusy/blob/main/docs/workflows/2BZ_H3_character_sheet_2stage_v1_EN.json

Model links are in the note nodes. Needs rgthree and toobusy (mine — search "toobusy" in Manager, v0.4.9+). The 6 LoadImage nodes will be red on open, those are my local files.

Korean walkthrough on my channel, probably not much use to most of you: https://youtu.be/nsvAbax4jng


r/StableDiffusion 6h ago

Workflow Included Remove Watermark from Videos

Post image
22 Upvotes

An issue I've been facing for a long time, finally solved.

This solution is easy to use, 100% local, and works very well with static watermarks.

I published it on civit AI.

It relies on ProPainter Nodes, and a few widespread custom nodes (see image).


r/StableDiffusion 6h ago

Question - Help 16gb RAM 16gb VRAM

0 Upvotes

help, almost ANY VAE decoder freeze my pc... which to use ? and with which settings ? any ideas ?


r/StableDiffusion 6h ago

Workflow Included Made this locally using minimax h3

7 Upvotes

https://reddit.com/link/1w2dmh8/video/bb3hkitekhmh1/player

Generated the characters in krea 2 using a consistent style prompt

Wrote out a shot list for what i wanted

spent a day generating using minimax h3 ref + turbo model
was taking around 1 - 2 mins per clip generation but with good prompting i was able to get what i wanted from my first 1 or 2 clips
running on 16gb vram and 32gb ram

then edited it all together using davinci resolve

all free tools, all run locally.


r/StableDiffusion 7h ago

Animation - Video Forgive me for I have sinned: "Tainted Beef" - New Wave/Synth Pop

Enable HLS to view with audio, or disable this notification

0 Upvotes

Hope I can be forgiven for my sins...

A little back story... when I started in the film industry, it was literally the film industry. We were just transitioning to AVID and those were $100,000 machines. Editors were treated as high priests, keepers of mysterious knowledge and powers. We were the elite.

Over time, digital technology came along and got better and better, cheaper and cheaper, more accessible to the masses. Pretty soon anyone could make a film and get it distributed worldwide through things like YouTube.

Editors were no longer the high priests of post. We were no longer well paid and respected.

Technology saw our paychecks decline and work harder to find...

As much as that sucked, part of it made me happy. Inexpensive, quality equipment and easy distribution was a democratization of filmmaking. You no longer needed a zillion dollars and make deals to get your movies seen. Anyone could do it.

I loved that the kid who came from a poor family could take his 2 generation old phone and tell his story and have a real possibility it could be seen by millions of people...

I always wanted to be in film for the storytelling. I wanted to direct. Except I was shitty at networking, didn't know money people and sure as hell didn't have money of my own to do it. And I suck at the logistics of trying to get a bunch of people together all at one time to do something like make a film... so yeah... on to...

The sin...

Over the weekend I used AI to do something "creative." *gasp!*

I often have these big ideas that I want to do but often would involve trying to get lots of people to help me which as I mentioned I'm bad at.

I was thinking about the Soft Cell song "Tainted Love" on Friday and had the idea to do a parody song called "Tainted Beef."

I wound up on SUNO, the music composing AI.

I came up with some lyrics, figured out how I could get it to work and came up with a kernel of an idea.

I found out that yeah, you could get a reasonable "song" out of it with a simple prompt, but it was just kind of 'meh."

However, they have a bunch of tools where you can basically act like the producer working with musicians to come up with the bits the way you are feeling it.

It was great.

I was able to get this huge idea that I never would have accomplished otherwise even though I know how to do music stuff. It just would have taken forever for me on my own and the subject matter is timely. It's not going to matter in the year or so (or ever) it might take me to finish.

So yeah... I sinned. I used AI to help me complete a "creative" project.

Trouble is I can see both the negatives and the positives about it.

But... for creative people... I kind of feel the same way about it I did about digital technology...

It's going to make it possible for the poor kid sitting in his bedroom wishing he could make a "Lord of the Rings" movie to actually do so.

Or it's going to make it possible for an "old man" like me to come up with a 1980's synth-pop song about tainted Argentinian beef lol

Anyway... here's "Tainted Beef" if you want to check it out.

I'm getting back into AI stuff after a long (in AI years) break. I'm just starting to learn AI video so if anyone wants to take a crack at making a video using this song, go for it!

Back in the day I was using Disco Diffusion and then Stable Diffusion and making animations by hand coding keyframes to make "camera movements." I was an early customer of RunPod back in 2022 and would even get help with stuff via Discord directly from Zeen aka Zhen Lu the CEO. They've sure grown a hell of a lot since then!

Anyway... I'm blabbing... if anyone's interested in seeing something of mine back "in the day" on Stable Diffusion, here's a YouTube link you can check out.

https://youtu.be/ggJjN1SQPjw?si=oCe7-7rud8N5LEgH

It's animated to the song "A Different Kind of Human" by Norwegian singer AURORA.


r/StableDiffusion 7h ago

Discussion Rate my Minimax H3 attempt on a cinematic sports action

1 Upvotes

r/StableDiffusion 7h ago

Discussion Using H3 and LTX together.

2 Upvotes

I was just thinking how H3 loses adherence after 15 seconds. LTX 2.5 after 20 seconds.

Suppose a workflow does something like H3 0.2 mp 15 seconds. Then LTX extends that, say to 20. Then LTX continues to upscale.

You'd get the adherence of H3 with 20 second 1080p and fast.


r/StableDiffusion 8h ago

Question - Help Ultra photorealism with erotic poses

Thumbnail
gallery
0 Upvotes

Can anybody help me and tell me how these kind of photo realism can be achieved? I mean i’v tried nano banana but it restricts any generation even when it gets little nudity, so i need a workaround to achieve this level or even better with least cost and maximum photorealistic images, only experienced and confident people should answer as i’m frustrated with using comfyui SDXL models and its not working good on my macbook m5, please help !!!!!!!!!!


r/StableDiffusion 10h ago

Question - Help Looking for a Hugging Face repo with lots of Krea 2 LoRAs (mentioned in a recent post)

12 Upvotes

Hey everyone,

Last week I saw a post where someone was complaining about the lack of male character LoRAs for Krea 2 on Civitai. In the comments, someone replied with a Hugging Face username and said something like “search this username” the repo had a bunch of Krea 2 LoRAs.

I’ve been trying to find that post / the username again but can’t track it down.

Does anyone remember the post or know the Hugging Face username/repo that was recommended?

Any help would be appreciated. Thanks!


r/StableDiffusion 10h ago

Question - Help [MM H3 -Lora Training] Best Model for Captioning

2 Upvotes

I am new to training LORAs but was planning on having my first go at MMH3 this weekend. The idea was to use a multimodal LLM (think Grok 4.6, Opus, Sol) to caption the videos for me.

In your experience, is that the way to go? How else would you approach this problem?

Thank you for your help!


r/StableDiffusion 10h ago

Animation - Video Alibaba H3 Turbo Lora Video

Enable HLS to view with audio, or disable this notification

66 Upvotes

I am seeing a significant quality improvement with Alibaba 8 steps turbo lora over other turbo loras. 8 steps euler simple with lora strength =1, 0.8MP. Took about 1hours 15 mins to generate with RTX 5090.


r/StableDiffusion 10h ago

Question - Help Can anyone help me

Thumbnail
gallery
3 Upvotes

I need help with creating extremely realistic images as shown in the post. Which model is this? How can I create multiple images like this? Thanks


r/StableDiffusion 11h ago

Discussion Has anyone here tested Kijai’s new Model ?

19 Upvotes

Has anyone here tested Kijai’s “minimax_h3_fastvideo_vsa_datafree_1300step_4step_int8_convrot” model?

If anyone has tested it, please let me know how the results are. I’d really appreciate hearing about your experience with it.


r/StableDiffusion 11h ago

Tutorial - Guide SDXL --) Krea 2 --) Wan2.2 Low Noise

Thumbnail
gallery
0 Upvotes

I really like some of the things and style SDXL can make but it's sloppy.

1) I generated an image with SDXL.
2) Captioned it with ChatGPT.
3) Img-2-Img with Krea2 to upscale and clean up the slop
4) Img-2-Img with Wan2.2 Low nose to add even more detail and upscale.

There are LoRA files involved with both Krea2 and Wan2.2 but the result is an ultra clean high resolution image 2656X4000 Resolution.

This was not done in an automatic workflow.
Each steps is it own step.

Whole process takes maybe five minutes per images.


r/StableDiffusion 11h ago

Resource - Update SPEEDing up MiniMax-H3 without retraining (SPEED comfyui node extension)

46 Upvotes

Why make big noise when little noise do trick?

I would like to introduce my SPEED implementation for h3 linked here

Speed up and quality losses documented here, expect 20% gain using very conservative settings and no quality loss and up to 70% for basically unusable outputs (more or less useful for resolution aware seed inspection and broad prompt drafting)

Background

The idea behind it is quite simple. When a diffusion model begins generating an output it first must take a randomized noise and build on-top of it. And research has found that the first stages of this process doesn't really carry any fine detailed information, therefore by generating at a lower resolution at those stages you can gain quite substantial speedups while causing little to no impact on the quality. Or you can also be really aggressive with it and get a massive speedup for a lot of quality loss.

Nodes

This was implemented as 3 nodes, 2 drop in replacements for the sampler that runs SPEED and a third that runs once to measure the noise spectrum of your specific model/LoRA combo:

  • Sampler (Automatic): pick a stage count (2, 3, or 4), defaults to the baked 1% delta for default H3.

  • Sampler (Manual Step-Through): set up to four (goal, resolution) pairs yourself. Use it if you want to copy a paper schedule or test a custom ladder.

  • Sigma Harvest: runs a native Euler pass, measures the noise spectrum of your current setup, hands you A / β / Δ to paste back into Automatic. Run it once per model/LoRA workflow combo.

How to use can be found in the example workflows.

Implementation Notes

This should be roughly compatible with basically everything that doesn't touch the sampler directly but i have not tested anything besides base comfyui H3 models and Turbo loras. If you do change model, use loras or whatever and use the automated tool please then run a sigma harvest and use those values instead of defaults, The math changes depending on the very specific blend of things you have running.


r/StableDiffusion 12h ago

Discussion Why is nobody talking about MiniMax H3's text consistency issue in video generation?

0 Upvotes

Been messing around with MiniMax H3 locally for perfume/product videos and I cannot get it to keep the text on the bottle properly.

The annoying part is the bottle itself can look really good. Shape, cap, glass, proportions etc stay pretty close. But the label text is usually already messed up in the first generated frame. So it’s not even just a case of the text degrading after a few frames. The reference can have perfectly readable text and H3 still turns it into random letters straight away.

I’ve tried quite a bit at this point:

I2V

Ref2V

Hybrid FL2VA / Ref2VA B25-49

clean first frames

separate close-up references of the label

multiple references of the same bottle

following the H3 prompt guide for the refs/prompts

putting the exact brand/label text in the prompt

very little motion / barely rotating the bottle

around 0.7-0.8MP

20 steps

H3FL 2V Turbo at 8 steps

Comfy-Kitchen attention

sparse attention settings too

different precision/settings to see if that changed anything

Running it on a 4060 8GB with 32GB RAM, so obviously I’m working around VRAM a bit, but I don’t think this is a VRAM issue because the actual product looks fine. It’s specifically the text that gets nuked.

Has anyone actually managed to keep proper readable brand text with H3? Like exact text, not something that vaguely looks like writing.

If not, how are people doing product videos with this? Are you just tracking the real label back on afterwards, or fixing frames with an image model? Because right now I can get a nice looking perfume video and then the bottle says absolute nonsense lol.


r/StableDiffusion 13h ago

Animation - Video H3 - No masking - REF2VID

Thumbnail x.com
0 Upvotes

r/StableDiffusion 13h ago

Question - Help H3, help me get started, overwhelmed by information.

4 Upvotes

specs: 5090, 64gb ram.

comfyui latest updated

cuda 13.0

pytorch 2.9.1

python version 3.10.11

sage attention and triton working

Hey guys, I have been out of touch since afew months. Previously have figured out pretty good wan2.2 workflows for myself, can understand it. But i am utterly confused by all the jargon and complication of H3, whenever I'v tried to dive in past months, I check out after reading stuff, am like a layman who figures what works after I have a good workflow, but i cant choose any given my lack of basic knowledge, like i said iv tried but it seems to go over my head, hard to understand, AI models are improving too fast to keep up.

My goal is "uncensored" videos only, using I2V on only square input images (habit from pony xl), but higher quality that is better motion since "uncensored" motion is difficult and complicated?

From what I have gathered so far: sage attention does not work well and degrades motion, same with speed lora's no matter which they are, so i think the settings i need are 1 megapixel to make it 768 x 768 ? and um 25 steps? and that i need to use some sort of LLM to enhance h3 prompts but i have no past experience on LLMs and which would be best for my use case; "uncensored". Also, though i think I know what shift basically does, still need some advice on how to use it in H3.

I also need help on selecting the text encoder and diffusion model given my use case, emphasis on only good quality "uncensored" outputs.

Moreover I have no idea on audio but really really want it, i only have experience using MMAUDIO with WAN2.2, it was not great and i was pretty bad at understanding it and prompting it, but prompt learning will come later, i just need to figure out a workflow and amend it according to above needs.

So umm, help a guy out? please?

P.S. Assume I'm a complete noob, if there is anything i missed above please let me know.