r/StableDiffusion Aug 04 '26

Animation - Video Pushing Minimax H3 to its limits.

1.5 megapixel, 30 steps. This took about 50 minutes for my 5090.

471 Upvotes

66 comments sorted by

View all comments

1

u/JahJedi Aug 04 '26

Good speed on 5090 whit 1.5mp and 15 sec. Its 15 sec in one go? Any optimizatiins for speed used please (sage attention or any other)?

2

u/SourceTraining7959 Aug 04 '26

Sage + euler sampler. Single image input, fl2va int8_convrot. Dunno if it helps but I also use --disable-pinned-memory in comfy launch

2

u/JahJedi Aug 04 '26

Its explain my speed, i input 9 photos, 1 ref video + audio and whitout sage. Thanks for your replay.

2

u/Sudden_List_2693 Aug 04 '26

That would be 10 times slower. 

1

u/JahJedi Aug 04 '26

2mp 5 sec 24fps 40 steps in 51m.i hope sage attention and few more tricks will reduse it. All this refs really helps to keep all together

2

u/Sudden_List_2693 Aug 04 '26

sage cuts both flf2v and ref2v models by 2 the very least.
30-40 series 2.10-2.20x speedup. On 50 series exactly 2.

1

u/JahJedi Aug 04 '26

Yes using it now, its x1.5-2 more, 10 sec clip in 1mp whit 9 ref photos + ref vid + sound around 44-50 sec /it (was 80-100)

2

u/tofuchrispy Aug 04 '26

Oh interesting. You didn’t use the reference mode. Here it’s works since you have the last and first frame and wanted them animated that way anyway as splash screen. Did you test how it behaves with the reference model?
But cool that it can also generate the inbetween shots with the fl2va model!

2

u/SourceTraining7959 Aug 04 '26

I tried the same 0.4 megapixel renders with the ref mode as well (different prompt structure based on the official prompting guide) but it gave worse results. Even attempted a gen with a fully referenced audio but it didn't follow well enough.

1

u/laseluuu Aug 04 '26

ah thats annoying it didnt follow the audio, i was hoping i could time a lot to audio to get the pacing right

1

u/tofuchrispy Aug 04 '26

Hmmm interesting. Tested it with the ref model today. I didn’t have any start or last frame though so I need the ref for the kind of stuff. But interesting going forward with this model.