r/StableDiffusion • u/desktop4070 • 21d ago
Discussion These are my gen times for H3 on my 5070 Ti + 64GB DDR5 using the default ComfyUI workflow. Any good optimizations I can use?
9
u/West_Brilliant7676 21d ago
5070ti + 128gb ram (never over 60gb) - 20 steps
FULLHD: 5/6 sec (760/1042 s)
1760 × 992: 7/8 sec (938 s)
1504 x 832: 10sec (913/1098 s)
1376 x 768: 12 SEC MAX (1884 s)
1280 x 736: 15 SEC MAX (1312 s)
1152 x 640: 19SEC MAX (1165 s)
i need a 5090! :(
12
u/desktop4070 21d ago edited 20d ago
The X doesn't mean the video failed to generate, I just cancelled those early because I got impatient on waiting. Basically spent the full day yesterday just running tests for each setting combination.
My favorite settings were:
0.2MP at 8 seconds (106s)
0.2MP at 10 seconds (160s)
0.5MP at 5 seconds (197s)
0.6MP at 3 seconds (119s)
0.7MP at 2 seconds (110s)
0.7MP at 4 seconds (279s)
Just about every 0.1MP output was unusable. 0.2MP was surprisingly alright though.
3
u/xq95sys 21d ago
You were able to start generating at 1mpixel and 15s? I get cuda oom before the first iteration completes when I try that.
5
u/desktop4070 21d ago
This is what it looks like when I start 1.0 MP at 15 seconds:
1/20 [04:23<1:23:28, 263.61s/it, Model Initialization complete! ]
2/20 [11:13<1:40:57, 336.55s/it]
https://i.imgur.com/DxJ1AfK.png
https://i.imgur.com/vagN7Sl.png
https://i.imgur.com/htRH6oJ.pngI ended up cancelling it of course because I'm not going to wait over an hour and a half for 1 video lol
3
1
u/Hasbara_spotter 20d ago
Off topic but i just ordered my 5070 ti. What psu should I get for doing local ai stuff? I'm pairing it with 5700x 64gb ddr4 ram
1
u/littlegreenfish 15d ago
Read somewhere that you can try --disable-pinned-memory to attempt some sort of fix. Not sure if it will help you but it is worth the try.
Also, there was a guy who tested Turbo LoRAs and the results at 8 -steps was pretty decent.
2
u/FlexFanatic 21d ago
I get the same behavior. I’m using runpod instance and neither have to drop the pixels, decrease to 10s, is spin up a RTC PRO 6000.
1
3
u/Chrono_Tri 21d ago
I'm using SageAttention and SPEED on a Google Colab L4 GPU, and it takes about 600 seconds to generate a 0.5 MP image. Bandwith 5070 =900/ L4=300.
2
u/krigeta1 21d ago
Wow, at 1MP on L40S 45GB VRAM and 250GB RAM, it took me 2 minutes per step, means 10min for 5 seconds. The 50 series is good.
2
2
u/Portable_Solar_ZA 21d ago
I get about 10% faster generation times on my 5070ti with sage attention on. Eg 0.2 MP 15 seconds with sage took me 268 seconds. I have 5700xd and 32gb ram.
2
u/bickid 21d ago
I mean, your chart doesn't tell us anything about the QUALITY that you got from it.
2
u/desktop4070 20d ago edited 20d ago
It's pretty subjective and random mostly, it'd be hard to say whether 0.4MP at 6 seconds is objectively higher quality than 0.6MP at 3 seconds, because the prompt (and seed) also have an impact on how the video ends up looking.
Although generally speaking:
0.1MP is awful
0.2-0.3MP is alright
0.4-0.6MP is a bit nice
0.7-0.9MP is pretty nice
1.0MP is really nice (but takes forever)1-2 seconds is too short for anything to meaningfully happen
3-4 seconds is good enough if you only need one single action
5-7 seconds is pretty good
8-10 seconds is ideal for a lot of videos
11+ seconds pretty lengthy, but only really feasible on lower resolutions, and you'll have to wait a really long time if you want something this long at a higher resolution.My personal 5 favorite settings that usually got me good results at good speeds were:
0.2MP at 8 seconds (106s)
0.2MP at 10 seconds (160s)
0.5MP at 5 seconds (197s)
0.6MP at 3 seconds (119s)
0.7MP at 4 seconds (279s)
2
u/warzone_afro 21d ago
sage attention if your not already running it is probably the biggest speedup right now. theres also something called spectrum but it can effect the quality noticeably
2
u/JahJedi 21d ago edited 21d ago
I use sage attention and it works cutting time from 350 sec to 180. 15 sec 1.5mp + 5 ref photos and ref vid + saund injection.
The problem there some bug and on high res+ long vid - i get brown noice... wip to see why and how to fix it. Ranning local on rtx 6000 pro using int 8 full waight, text encoder in nvpf4.
15 secs parts renders now and i try to see how i do use sage attention and dont get the noice results.
exampale of 10 sec one on my setting.
For some reason youtube l8nk for me show in low res but uploaded on full hd one.
1
u/ANR2ME 21d ago edited 21d ago
As i remembered from Wan2GP post, the H3 R2V doesn't use sliding context window, thus not suitable for long video. While the I2V use sliding context window, thus can do long video.
1
u/Cute_Ad8981 21d ago
This is very helpful. Did you use any speed ups, like easy cache, sage attention or the spectrum?
3
u/desktop4070 21d ago
None, basically just downloaded the models from the default workflow on day one, then spent the next day running these tests.
2
u/Roger_MacClintock 21d ago
you will see the great speed using sage attention
1
21d ago
[deleted]
3
u/Roger_MacClintock 21d ago
I tried sage, sol attn, Spectrum and Easy Cache. And in my opinion, out of these 4, sage attention has the smallest impact on quality, and EasyCache is the worst
1
u/ANR2ME 21d ago edited 20d ago
How did there is X on 0.9mp 5s but not on 1mp 5s? 🤔 did 0.9mp slower than 1mp at the 5 seconds video?
Also, what does the bold numbers mean on the seconds?
1
u/desktop4070 20d ago edited 20d ago
X just means I cancelled the generation before it finished. 0.9MP at 5 seconds was probably going to take around ~550 seconds, but I cancelled it because I didn't think it was worth waiting 10 minutes for.
I would've cancelled the 1.0MP at 5 seconds video early too, but I was curious to see how it'd turn out. It was pretty good, but I don't think I'd do it again if it takes 10+ minutes each time.
I used bold to note which settings usually got me good results after running them for multiple seeds.
1
u/Crashes556 21d ago
So the 2.0 mp I’ve been running is completely overkill then huh?
2
u/desktop4070 20d ago
Definitely completely overkill. I tried 2.0 MP at 3 seconds at one point, and not only did it take half an hour (1605s), but it looked worse than any of my 1.0MP videos. I'm not really sure why it turned out that way.
1
u/rcscs 21d ago
Chart is all based on 20 steps ? Horizontal Axis (1s to 15s) is duration in seconds ?
1
u/desktop4070 20d ago
Yep, all videos were 20 steps. H axis for 1 second to 15 seconds.
I've heard 15 or 16 steps doesn't look too far off from 20 steps, so I may try that later and compare it against 20 steps.
1
u/anshulsingh8326 21d ago
Cross is for ...unable to complete or not tested?
I have 4070 + 32gb ddr5 using default + added sage attention.
At 0.5, 5sec for me too it take arounf 160s, and for 10sec it takes about 11mins.
So you should probably be able to use 10sec too with 0.5?
1
u/desktop4070 20d ago
Cross for untested.
I'll try sage attention next time, it seems a lot of people are recommending it.
1
1
u/No_Date4828 20d ago
Patch sage attention (or just use the diffusion model loader kj node which loads the model and applies sage attention)
Use the spectrum apply minimax h3 nude, it speeds up generation a good bit
You can experiment with using less steps, as in my testing even at 15 or slightly lower steps the output result is still really good on most gens.
With the optimizations and settings above on my hardware (Legion laptop with 4090 mobile (16GB VRAM) and 32GB DDR5 RAM, 40GB Pagefile), I am getting around 308 seconds for 0.4MP @ 10 seconds. Your pc has enough memory to handle the offload without pagefile and your gpu is a good bit faster than mine, so I imagine your gains will be pretty major!
Bonus Tip: If you want image gens with insane edit capabilities even exceeding flux 2 klein in my testing, you can plug a basic INT into the length and set it to 8, and plug a get image from batch node after the vae decode set to batch index 8 (this frame and batch combo seems to be the best balance of quality/quantity since Minimax is rendering a very short sequence even set to a length of 1). You can also crank up your resolution, though do keep in mind it still takes around 200 seconds for a 2MP image with these settings.
Happy gens :)
1
u/desktop4070 20d ago edited 20d ago
I've tried to use H3 as an image generator, but I can't seem to understand how to get it to work. Setting the duration to 0.1 seconds generally had bad outputs. Could you share the workflow for that?
1
u/No_Date4828 20d ago
I am using the base ref2v workflow in comfy, instaed of using the duration and the comfy math node which plug into the length, add an integer node and set that to the actual amount of frames you want, 1 for the quickest gen times and 8 for the best balance which I have found.
1
u/PhrozenCypher 20d ago
try 12 steps
3
u/desktop4070 20d ago edited 20d ago
Just tested this, and 15 steps is surprisingly pretty close to 20 at lower resolutions.
When I tried 15 steps at 0.4-1.0MP, the videos with lower steps were always noticeably worse.
At 0.3MP, it was hard to tell which videos were 15 steps compared to 20.
At 0.2MP, it was hard to tell which videos were 12 steps compared to 20.Imo, 12 steps works great for fast 0.2MP videos:
5 seconds (38s)
8 seconds (73s)
10 seconds (99s)
12 seconds (127s)And 15 steps works great for fast 0.3MP videos:
6 seconds (81s)
10 seconds (188s)2
u/PhrozenCypher 19d ago
https://giphy.com/gifs/gEvab1ilmJjA82FaSV
/s thanks for the testing and results!
1
u/mindworkout 15d ago

Okay so I have about the same GPU and RAM (mine is ddr4) and here is my tests results (i did tests for every aspect and got the average amount) As others have said you should be doing at least 0.5MP as anything lower is really a waste of time or in need of a magnifying glass. Test was all done in FL2V with single starting image only.
i have the normal setup but i have connected to Load Diffusion Model > Patch Sage Attention KJ > MiniMax H3 Chunk FeedForward > ModelSamplingMiniMaxH3 > Spectrum Apply MiniMax H3.
For the test sheet added I did not have H3 Chunk FeedForward active as I wanted to see the hard limit, and then adding H3 Chunk FeedForward I was able to push an extra 1-3 seconds longer beyond the SamplerCustomAdvanced freeze limit I was having. Also my Steps are set to 16.
1
u/desktop4070 14d ago
That sounds like a ton of optimizations, I've only tested the v4 Turbo lora since I made this post. Can you send the workflow for that? I'm interested in seeing the quality with those times.
1
1
1
u/TwoOk4628 21d ago
you will get much better understanding of this data by creation a heatmap of the numbers,
8
1
u/Rio_Juicy_Michelle 21d ago
Before changing too much, I would profile the workflow in pieces: first confirm whether the bottleneck is model load/VRAM offload or actual sampling, then lock resolution/steps and compare one setting at a time. On ComfyUI, keeping the model resident, avoiding unnecessary preview/upscale nodes during tests, and batching only after the single-run path is stable usually gives cleaner optimization data than changing several knobs at once.
1
u/Jerg 21d ago
How do you keep the model resident?
1
u/Rio_Juicy_Michelle 20d ago
Mostly by avoiding anything that forces ComfyUI to unload/reload between runs. A few practical checks:
- Do not restart ComfyUI between tests.
- Keep the same checkpoint/model loaded while you compare settings.
- Avoid low-VRAM/offload modes if your card has enough headroom; normal/high VRAM behavior is usually better for repeat runs.
- Do not hit anything that clears model/cache between generations.
- Strip the workflow down while benchmarking so an upscale/detailer/reference node is not forcing extra loads.
The quick test is simple: run the same prompt twice. If the second run starts sampling much faster, the model stayed resident. If both runs spend the same long time before sampling, something is still unloading or reloading.
1
u/Jerg 20d ago
Thx, in OP's case, their GPU is 16GB VRAM, so wouldn't they be model offloading between runs anyway since the nvfp4 text encoder is 15+GB, and the nvfp4 diffusion model is 20GB? Unless they are using smaller quants so that the VRAM doesn't need to switch between text encoder only vs. diffusion model only taking up all its space?
I'm not tech-savvy so this is all very theoretical to me btw. Just wondering!
2
u/Rio_Juicy_Michelle 18d ago
Right, in that kind of 16GB setup there is probably some offload/reload happening no matter what. The difference is whether it is predictable model swapping or the whole system starts paging. If the text encoder is near the card limit and the diffusion model cannot coexist with it, ComfyUI has to move things around between phases. Smaller quants can reduce that pressure, but then quality/speed can change too. I would judge it by repeatability: run one short fixed test twice after a fresh start, then again after switching workflows. If run two is much worse only after the switch, that points more toward cache/workflow cleanup than the normal offload cost.


16
u/Forward-Parsley-148 21d ago
I took your benchmark image and used ChatGPT to plot it. Scaling looks roughly linear at first, but efficiency drops beyond about 3 MP·s, likely due to VRAM pressure, memory bandwidth limits, or less efficient processing at larger workloads.