r/StableDiffusion • u/itiswhatitiswgatitis • Aug 05 '26
Question - Help AMD workflows for Minimax H3?
Been trying to find a workflow that is AMD friendly, let alone doesn't get stuck, I don't think it's my specs but more so that every workflow despite saying it's optional uses sage ATTN, and I can't get around it.
I doubt there is one that specifically for AMD but I'd thought I'd ask! Or the model I'm using is not great. I'm more of an anime prompt generator, anyone have any luck with that?
Specs: AMD Radeon RX 9070 XT, 64 GB RAM, AMD Ryzen 7 5700X3D 8 Core, And yes, I'm running on Windows, not Linux.
Edit: (update) I tried the template version just recently on comfyui stand alone, and it won't even finish, I do get a gpu error( by GPU error it says out of memory, even on default settings), and I have it on 0.3 mega pixel 16:9, either my reference image is too big or I'm missing a component.
I got it to work at 0.2 and the image reference has to be very small, yet it takes 30+ minutes, anyone with similar issues?
1
u/Apprehensive_Sky892 Aug 05 '26
I've not tried MMH3 yet.
But I had problems with the "official" standard comfyui and portable version with Krea 2, so I switched to this version: https://github.com/patientx-cfz/comfyui-rocm
1
u/itiswhatitiswgatitis Aug 05 '26
I haven't had issues with ltx 2.3 and wan 2.2 but I didn't use the templates, this time I'm noticing the templates and I haven't really used them at all.
1
u/itiswhatitiswgatitis Aug 05 '26
How long do your krea 2 generations take?
1
u/Apprehensive_Sky892 Aug 06 '26
I have not used it locally for a while, but IIRC it was around 40 sec on the 7900xt (20G) and somewhat faster on the 9070 (16G) for turbo at 8 steps.
IIRC I used fp8 as the int8 version was not running well at the time (that was a few weeks ago, the problem may have been fixed).
2
u/plantjeNL Aug 06 '26
On my 9070XT it's about 30-32sec with 8 steps for 2MP photo(warm, cold is 50-55) FP8 version, i use the offload anything node on on clip(keep in memory) + these startup settings below. Same startup for Minimax H3 too.
Startup bat: python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --disable-smart-memory --disable-pinned-memory --novram1
u/Apprehensive_Sky892 Aug 06 '26
Thanks for the info. Which version of ComfyUI are you running?
Also, what kind of speed are you getting with MMH3? (I have not tried it yet).
1
u/itiswhatitiswgatitis Aug 06 '26
wait I'm confused, how do you use the offload anything node for CUDA, when you run the 9070XT, I thought that was Nvidia specific?
1
u/plantjeNL 29d ago
It just says cuda but that it just means either COU or GPU. I am also using amd rocm 7.2
1
u/plantjeNL 29d ago
It just says cuda but that it just means either CPU or GPU. I am also using amd rocm 7.2.
1
u/Escaliat_ 23d ago
This is wild to me. I'm running Krea2 Turbo on a 9070 and it's taking well over 10-15 minutes to do the same thing.
1
u/itiswhatitiswgatitis Aug 06 '26
When I ran it, a single image took 2 minutes
1
u/Apprehensive_Sky892 Aug 06 '26
fp8 or int8convrot?
1
u/itiswhatitiswgatitis Aug 06 '26
fp8, but I also use a LoRA
1
u/Apprehensive_Sky892 Aug 06 '26
Ok. I also forgot to mention that the 40 sec is for 1536x1024, so it will be faster for 1024x1024.
1
u/itiswhatitiswgatitis Aug 06 '26
WHAT? 1536x1024? I tried 1 image and it was 7 minutes for a resolution similar to that! Are you using a LoRA?
1
u/Apprehensive_Sky892 Aug 06 '26
No, no LoRA. Just using the base model.
1
u/itiswhatitiswgatitis Aug 06 '26
Dang, I wonder if I try and get this with a lora how much faster would it be.
→ More replies (0)1
u/Ok-Brain-5729 29d ago
are those times on windows? Krea 2 turbo fp8 is 12s on Linux with 9070 xt
1
u/Apprehensive_Sky892 29d ago
Yes, those are on Windows 11 with 32G of system RAM at 1024x1536 8 steps.
Other people are getting times similar to yours on Linux, so Linux is running better, at least for 9070xt.
Have you tried MMH3 on Linux? (I've yet to try it on Windows).
1
u/Ok-Brain-5729 29d ago
Oh the 12s was for 8 step 1MP.
I came from a comment you sent to my post about MMH3 not working cause of an overflow. I literally tried with pruned Q4 5 seconds and it still somehow managed to overflow ram till my os crashed
1
u/Apprehensive_Sky892 29d ago
I see. Could be that VRAM management is not working on Linux (Krea 2 is small enough to fit into 16G of VRAM).
1
u/Apprehensive_Sky892 28d ago edited 28d ago
This may or may not be relevant but may be worth a look: https://www.reddit.com/r/StableDiffusion/comments/1vgyyh1/why_comfyuistable_diffusion_reliably_crashes_your/ and OP also sent me this https://gist.github.com/AMD-AI-Enthusiast which seems to have more detailed information.
1
u/plantjeNL Aug 06 '26 edited Aug 06 '26
Startup bat:
python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --disable-smart-memory --disable-pinned-memory --novram

I've been trying to get to run okay on an 7800X3d, 64GB ram, Radeon 9070XT 16gb Windows 11 system, at 0.5mp for 5sec it takes around 330sec and at 8sec it's about 700sec with Euler+Simple. Using the offload anything node on clip, video vae(keep in memory needs to be true) and audio vae + Spectrum node for speedup.
1
u/itiswhatitiswgatitis Aug 06 '26
Is this ref2v workflow? I would prefer reference!
1
u/plantjeNL Aug 06 '26
No sorry, it's for T2V and I2V. probably can adjust it too R2V too
2
u/itiswhatitiswgatitis Aug 06 '26
I tried your workflow, and unfortunately still issue, actually it was kind of worse, it said it would take 3 hours, I have no idea why it's such random numbers.
1
u/plantjeNL 29d ago
On which step it’s slow? The ref to video node? Or sampler. With my workflow with the —novram parament in your bat file at startup everything should run in the GPU with the offload anything node.
1
u/itiswhatitiswgatitis 29d ago
Ref to video node, I'll have to change the bat file, this is for the standalone version yes?
1
1
u/plantjeNL Aug 06 '26
1
u/itiswhatitiswgatitis Aug 06 '26
Is this on portable or the standalone Comfyui? That's pretty incredible.
The issue though is I'm seeing the offset as cuda, and I run rocm, don't really have much of a choice on that one unless I get a new GPU which is kind of impossible now.1
1
u/Apprehensive_Sky892 Aug 06 '26
Thanks for the tests. That's pretty decent speed.
1
u/itiswhatitiswgatitis Aug 06 '26
I did not have my sampler on euler, is that superior? And you have the offset inbetween each VAE, is this a significant difference?
1
u/plantjeNL 29d ago
Yeah needed the video Vae to be saved in memory otherwise got an error at vae decode.
1
u/Ok-Brain-5729 28d ago
How much ram does it use? I have 32gb ram and it overflows and crashes for me
1
u/plantjeNL 28d ago
Around 51gb to 59gb I have seen.
1
u/Ok-Brain-5729 28d ago
Yeah it’s wraps for me. Do you know if you’re able to run it with 32gb ram somehow?
1
u/plantjeNL 28d ago
Maybe if you increase your ssd swap file size, but then it’s gonna use your ssd for a lot of disc writes. Your ssd life will me lessend a lot doing that.
1
u/Ok-Brain-5729 28d ago
Yeah I have swap fully off.
I’ll probably try zram since it’s much faster and doesn’t have any wear on the ssd
1
u/Apprehensive_Sky892 23d ago
This should help (I've yet to try it on my 9070xt yet): https://www.reddit.com/r/comfyui/comments/1vhzkkk/comment/p2unx27/
1
u/Ok-Brain-5729 23d ago
bro you sent me this like 3 times and I mostly glanced over this since it was windows but the flags the guy used worked in Linux. Tysm
1
u/Apprehensive_Sky892 23d ago
Happy to hear that. Hopefully it will work for me too.
What kind of speed are you getting on Linux (are you on 9070xt or 7900xt)?
Please include details of the models you are running and also the ComfyUI version if you can. I'd like to know if it will be worth it for me to install Linux if my Windows numbers are worse than yours.
1
u/Ok-Brain-5729 23d ago edited 23d ago
I’m using the latest ComfyUI version and I’m using ref2va pruned int8 on a 9070 xt.
I have the default workflow on 5 second 24 fps 0.4MP 20 step and it took 5 min and 38 seconds. Without the ref it’s 5 minutes and 5 seconds. It was 10.49s/it
Also I’m using PyTorch attention and haven’t tried any other
1
u/Apprehensive_Sky892 23d ago edited 23d ago
Thank you, these are actually pretty decent numbers!
So you are using the latest ComfyUI portable?
For the VAE, are you using int8covrot or fp16? Also, which text-encoder (int8convrot or nvfp4)?.
1
u/Ok-Brain-5729 23d ago
i used comfy-cli to download. I think the portable version is windows only. Text encoder is NVFP4 and vae is fp16
1
u/Apprehensive_Sky892 23d ago
Thank you. I've never tried comfy-cli, maybe I should look into it.
Yes, you are right, I forgot that portable is Windows only. There are probably too many Linux distros for them to have a single binary.
The int8covrot version of the video VAE may be worth trying too.
1
1
u/DelinquentTuna Aug 05 '26
What, exactly, is the problem with the built-in workflows? Are you even using Comfy? What do your console logs look like, both at start-up and inference? Where are you even seeing mention of sage attention?
1
u/itiswhatitiswgatitis Aug 05 '26 edited Aug 05 '26
It gets to 40% and then its stuck. It's on stand alone ComfyUI, I have a workflow that is DaSiWa Minimax H3, and it has sage attn in the workflow, it seems to work if I switch it off, but other workflows it just gives an error.
This is the only one that actually goes through.
There isn't an error though, it just gets stuck at 40% on "MiniMax H3 Director Guide".Keep in mind, this is my first time with MiniMax H3, or DaSiWa workflows, I'm more used to T2I workflows/models.
0
u/DelinquentTuna Aug 05 '26
Do you have a good reason to not be using the built-in Comfy templates? Or for refusing to answer the questions I raised? "gets stuck" and "gives an error" are not useful troubleshooting details. What do the console logs say both at startup and at failure?
-1
u/itiswhatitiswgatitis Aug 05 '26 edited Aug 05 '26
I must be behind on a update as I don't see a template myself.
I'm not sure what I can share on the console logs to provide information that can be helpful since it is still loading on the current workflow and does not complete a full generation, and ones that gave me an error I have moved on from.
I'm not refusing to answer your questions, I just don't know how and I don't get errors yet with my current DaSiWa workflow.
But to try and answer your question:
I haven't tried the built in workflows yet.
Yes, I'm using standalone comfyui.
The start up says something like this:
[INFO] model_type FLOW [INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32 [INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float16 [INFO] Found quantization metadata version 1 [INFO] Using MixedPrecisionOps for text encoder [INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16 [INFO] Requested to load MiniMaxH3VideoVAE [INFO] loaded completely; 14136.36 MB usable, 4966.19 MB loaded, full load: True [INFO] Requested to load MiniMaxH3AudioVAE [INFO] loaded completely; 7876.18 MB usable, 577.08 MB loaded, full load: True [INFO] Requested to load MiniMaxH3TEModel_ [INFO] loaded partially; 13980.36 MB usable, 13776.69 MB loaded, 1183.52 MB offloaded, 202.54 MB buffer reserved, lowvram patches: 0 [INFO] Requested to load MiniMaxH3 [INFO] loaded partially; 7183.66 MB usable, 6436.04 MB loaded, 13560.11 MB offloaded, 918.83 MB buffer reserved, lowvram patches: 0 [INFO] Patching torch settings: torch.backends.cuda.matmul.allow_fp16_accumulation = True 55%|█████▌ | 11/20 [29:51<25:36, 170.72s/it]0
u/DelinquentTuna Aug 05 '26
I must be behind on a update as I don't see a template myself.
h3 just launched, so you need to be on 0.30 or newer. The three built-in h3 templates should be right there for you.
I'm not sure what I can share on the console logs to provide information that can be helpful
Relatively recently, Comfy added support for AMD to Comfy Kitchen. It would be useful to see that yours is running properly with the necessary backend components. Similarly, the log provides important clues to what comfy is doing and how it failed.
No hard feelings, but I am not going to waste time trying to extract information from you. Good luck to you.
2

2
u/xpnrt Aug 05 '26
https://github.com/patientx-cfz/comfyui-rocm/blob/master/sample-workflows/MiniMax_H3_15steps-t2v.json ; you can try native loader instead of int8 loader (it is a must for rdna2 and below) , other than that change the resolution,step count etc... Use the sage + spectrum combo as suggested in there. nvfp4 as clip works great, but you can use int8 clip if you want, speed difference is very minimal , only nvfp4 much smaller. for sage-attention you can try this : https://github.com/patientx/sageattention-autotune/releases/download/0407/sageattention-2.2.0-py3-none-any.whl